Methods and processes for evaluating gene mutations
The method and system improve the detection of small CNVs in genomic regions by using genome-wide and intensive sequencing analysis with segmentation and z-score cutoffs, addressing the limitations of conventional methods and enhancing diagnostic precision.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SEQUENOM INC
- Filing Date
- 2025-02-14
- Publication Date
- 2026-07-29
AI Technical Summary
Existing methods struggle to accurately detect small copy number variations (CNVs) in genomic regions, such as microdeletions and microduplications, which are often undetectable by conventional cytogenetic methods, hindering precise genetic diagnosis and predisposition analysis.
A method and system for classifying the presence or absence of CNVs in subchromosomal regions using genome-wide and intensive sequencing analysis, involving segmentation processes and quantitative value determination of sequence reads, optimized for precision using training sets and z-score cutoffs.
Enhances the detection accuracy of microdeletions and microduplications, enabling precise genetic diagnosis and predisposition analysis by identifying CNVs with high sensitivity and specificity.
Smart Images

Figure 0007897354000003 
Figure 0007897354000004 
Figure 0007897354000005
Abstract
Description
[Technical Field]
[0001] Related patent applications This application claims the benefits of U.S. Provisional Patent Application No. 62 / 449,766, filed on 24 January 2017. The entire contents of the said Provisional Patent Application are incorporated herein by reference for all purposes.
[0002] field The technologies provided herein relate in part to methods, systems, instruments, and computer program products for the non-invasive classification of gene copy number variations (CNVs) in test samples. The technologies provided herein are useful, for example, for the classification of gene CNVs in samples as part of non-invasive prenatal testing (NIPT) and oncological testing. [Background technology]
[0003] background The genetic information of living organisms (e.g., animals, plants, and microorganisms) and other forms of genetic replication (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a sequence of nucleotides or modified nucleotides that correspond to the primary structure of chemical nucleic acids or hypothetical nucleic acids. The human genome contains approximately 30,000 genes located on 24 chromosomes (i.e., 22 autosomes, the X chromosome, and the Y chromosome; see The Human Genome, T. Strachan, BIOS Scientific Publishers, 1992). Each gene encodes a specific protein, which, after expression via transcription and translation, performs a specific biochemical function in living cells.
[0004] Many medical conditions are caused by one or more gene mutations and / or genetic alterations. Certain gene mutations and / or genetic alterations cause medical conditions such as hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, D.N. Cooper and M. Krawczak, BIOS Publishers, 1993). Such genetic disorders can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a particular gene. Certain birth defects are caused by chromosomal abnormalities also known as aneuploidy, such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), monosomy X (Turner syndrome), and certain sex chromosome aneuploidy, such as Klinefelter syndrome (XXY). Another gene mutation is the sex of a fetus, which can often be determined based on the sex chromosomes X and Y. Some gene mutations can cause or lead to any of several diseases (e.g., diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancers (e.g., colorectal cancer, breast cancer, ovarian cancer, lung cancer, bladder cancer, stomach cancer, cervical cancer, kidney cancer, prostate cancer, brain cancer, and esophageal cancer)) in an individual.
[0005] Identifying one or more gene mutations and / or genetic alterations (e.g., copy number changes, copy number variations, single nucleotide changes, single nucleotide variations, chromosomal changes, translocations, deletions, insertions, etc.) or genetic variances can allow for the diagnosis of specific medical conditions or the determination of predisposition to specific medical conditions. Identifying genetic variances can facilitate medical decisions and / or enable the use of beneficial medical procedures. In certain embodiments, the identification of one or more gene mutations and / or genetic alterations requires analysis of circulating cell-free nucleic acids. Circulating cell-free nucleic acids (CCF-NA), e.g., cell-free DNA (CCF-DNA), consists of DNA fragments derived from cell death that circulate in the peripheral blood. High concentrations of CF-DNA can suggest certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infections, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and can be used in various non-invasive prenatal diagnostics. [Overview of the Initiative] [Means for solving the problem]
[0006] Abstract A method for classifying the presence or absence of copy number variations in a subchromosomal region of a test sample is provided herein, the method comprising the steps of: a) identifying the presence or absence of copy number variation segments in a region comprising a first set of genomic parts using a method comprising a segmentation process; and b) providing quantitative values of sequence reads for a subregion within a subchromosomal region comprising a second set of genomic parts, wherein the second set is a predetermined set of genomic parts, and the genomic parts in (a) and (b) comprise a portion of a reference genome on which the obtained sequence reads are mapped to nucleic acids in the test sample, thereby providing a classification of the presence or absence of copy number variations in a subchromosomal region of a test sample according to (a), or (b), or (a) and (b).
[0007] In a particular embodiment, a method is provided for classifying the presence or absence of copy number variations in a subchromosomal region of a test sample, the method comprising: a) identifying the presence or absence of copy number variation segments in a region comprising a first set of genomic subregions using a method comprising a segmentation process; b) providing quantitative values of sequence reads for a subregion within a subchromosomal region comprising a second set of genomic subregions, wherein the second set is a predetermined set of genomic subregions, and the genomic subregions in (a) and (b) comprise a portion of a reference genome on which the obtained sequence reads are mapped to nucleic acids in the test sample; and c) providing a classification of the presence or absence of copy number variations in a subchromosomal region of a test sample based on changes within the region in (a), within the subregion in (b), or both, relative to a set of reference samples. The region in (a) may encompass a subchromosomal region or may overlap with a subchromosomal region.
[0008] In some embodiments, the first set of genomic segments is a region within a chromosome where copy number variations associated with the desired phenotype are expected to exist. Such genomic segments can often be obtained by mining public disease databases, such as the International Standards of Cytogenomic Arrays database (ISCA). In one embodiment, the phenotype is a microdeletion syndrome. In one embodiment, the first set of genomic segments is one or more genomic segments selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23.
[0009] In certain embodiments, a method for classifying the presence or absence of copy number variations in a sub-chromosomal region for a test sample is provided. The method includes: a) providing a quantitative value of sequence reads for a sub-region within a sub-chromosomal region that includes a subset of genomic portions, where: i) the genomic portion includes the portion of the reference genome to which sequence reads obtained for nucleic acids in the test sample are mapped; ii) the set is a predetermined subset of genomic portions; iii) the predetermined subset of genomic portions is identified by a process that includes: 1) providing a plurality of candidate sub-regions within the sub-chromosomal region; 2) providing one or more accuracy metrics for each of the plurality of candidate sub-regions for a plurality of samples in a training set, where each of the plurality of samples is classified as having a copy number variation in the sub-chromosomal region; and 3) identifying the sub-region in (a) as a sub-region that provides an accuracy metric equal to or exceeding a predetermined threshold; and b) providing a classification of the presence or absence of a copy number variation in the sub-chromosomal region for the test sample according to the quantitative value of the sequence reads in (a) relative to the quantitative values of the sequence reads for a reference sample set.
[0010] A system including one or more processors and a memory is also provided herein. The memory includes instructions executable by the one or more processors, and the instructions executable by the one or more processors include a) configured to identify the presence or absence of copy number variant segments in a region including a first subset of genomic portions using a method including a segmentation process; and / or
[0011] b) configured to provide a quantitative value of sequence reads for a sub-region within a sub-chromosomal region containing a second genomic subset, where the second set is a predetermined genomic subset, and the genomic parts in (a) and (b) include parts of the reference genome to which sequence reads obtained for the nucleic acid in the test sample are mapped; c) configured to provide a classification of the presence or absence of copy number variations in the sub-chromosomal region for the test sample based on changes within the region of (a), within the sub-region of (b), or both, relative to a reference sample set.
[0012] A computer program product as a computer-readable storage medium is also provided herein, the product comprising instructions programmed for a computer to: a) identify the presence or absence of copy number variant segments in a region containing a first genomic subset using a method including a segmentation process; and / or b) instructions programmed to provide a quantitative value of sequence reads for a sub-region within a sub-chromosomal region containing a second genomic subset, where the second set is a predetermined genomic subset, and the genomic parts in (a) and (b) include parts of the reference genome to which sequence reads obtained for the nucleic acid in the test sample are mapped; and c) instructions programmed to provide a classification of the presence or absence of copy number variations in the sub-chromosomal region for the test sample based on changes within the region of (a), within the sub-region of (b), or both, relative to a reference sample set.
[0013] Certain embodiments are further described in the following description, examples, claims and drawings.
[0014] The drawings illustrate certain embodiments of the technology and are not limiting. For clarity and simplicity of illustration, the drawings are not drawn to scale and in some cases various aspects may be exaggerated or enlarged to facilitate understanding of certain embodiments. [Brief explanation of the drawing]
[0015] [Figure 1] Figure 1 shows an illustrative embodiment of a system capable of performing a particular embodiment of the present technology.
[0016] [Figure 2] Figure 2 shows the sensitivity values for detecting the range of microdeletions relative to fetal proportions for two detection methods (i.e., genome-wide sequencing and focused sequence analysis).
[0017] [Figure 3] Figure 3 shows the chromosome 22q11.2 deletion region associated with DiGeorge syndrome. The analysis of certain 22q11.2 deletions discussed herein includes the region indicated by the dashed vertical line.
[0018] [Figure 4] Figure 4 shows chromosomal 22q11.2 deletions present in genomic DNA (gDNA) reported in the ISCA database and used in mixed models. The black vertical dashed lines (i.e., the outer set of vertical dashed lines) represent the analysis windows for 22q11.2 deletions using the genome-wide analysis algorithms discussed herein. The gray vertical dashed lines (i.e., the inner set of vertical dashed lines) represent the focused analysis windows for 22q11.2 deletion analysis optimized around specific 22q11.2 deletion regions.
[0019] [Figure 5]Figures 5A–5D show schematic depictions of 22q11.2 deletions analyzed by whole-genome sequencing. The 22q11.2 deletion is represented by the simulated signal, noise, and event size. Panels A–D represent samples with possible combinations of lower or higher fetal proportions and smaller or larger event sizes. Figure 5A shows whole-genome sequencing analysis of a large deletion event in a sample with a low fetal proportion. The 22q11.2 deletion is represented by the simulated signal, noise, and event size. Figure 5B shows whole-genome sequencing analysis of a large deletion event in a sample with a high fetal proportion. The 22q11.2 deletion is represented by the simulated signal, noise, and event size. Figure 5C shows whole-genome sequencing analysis of a small deletion event in a sample with a low fetal proportion. The 22q11.2 deletion is represented by the simulated signal, noise, and event size. Figure 5D shows whole-genome sequencing analysis of a small deletion event in a sample with a high fetal proportion. The 22q11.2 deletion is represented by showing the simulated signal, noise, and event size.
[0020] [Figure 6] Figure 6 shows the sensitivity of the combined analysis for detecting the 22q11.2 deletion. [Modes for carrying out the invention]
[0021] Detailed explanation A useful method for classifying the presence or absence of copy number variations in a subchromosomal region of a test sample is provided herein. In some embodiments, the sample nucleic acid subjected to the sequencing process and the resulting sequence reads are further analyzed to determine the presence or absence of copy number variations. In some embodiments, the presence or absence of copy number variations is classified according to genome-wide sequencing analysis. In some embodiments, the presence or absence of copy number variations is classified according to intensive sequencing analysis (e.g., analysis of sequence reads for a given genomic subregion). Intensive sequencing analysis can improve the accuracy (e.g., sensitivity) for detecting copy number variations in a particular type of sample. In some embodiments, the presence or absence of copy number variations is classified according to genome-wide sequencing analysis and intensive sequencing analysis.
[0022] In some embodiments, systems, apparatus, and computer program products that perform the methods or parts of the methods described herein are also provided. Classification of copy number variations using genome-wide and / or intensive sequence analysis.
[0023] Methods and processes for classifying the presence or absence of copy number variations (e.g., microdeletions, microduplications) in subchromosomal regions are provided herein. As used herein, microdeletions and microduplications commonly refer to deletions or duplications smaller than 5 million base pairs. Microdeletions and microduplications are usually too small to be detected by conventional cytogenetic methods or high-resolution karyotype analysis. The methods and systems of this disclosure enable the accurate detection of both microdeletions and microduplications.
[0024] In some embodiments, the presence or absence of copy number variations is classified according to a set of sequence reads. In some embodiments, sequence reads are obtained for nucleic acids in a test sample. In some embodiments, sequence reads are mapped to genomic portions in a reference genome. In some embodiments, the classification of the presence or absence of copy number variations in a subchromosomal region includes identifying the presence or absence of copy number variation segments. As used herein, a copy number variation segment is a segment in a chromosome containing copy number variations. In some embodiments, copy number variation segments are identified using a method that includes a segmentation process. A method that includes a segmentation process may include a decision analysis, such as a decision analysis described herein. A method that includes a segmentation process may be part of a genome-wide sequencing method. A method that includes a segmentation process may be part of a sequencing analysis of nucleic acids captured by probe oligonucleotides. In some embodiments, the classification of the presence or absence of copy number variations in a subchromosomal region includes providing quantitative values of sequence reads for subregions within the subchromosomal region. As an illustrative example, a subregion is a region defined by the gray dashed line in Figure 4.
[0025] In some embodiments, a subregion comprises a predetermined set of genomic parts. Providing quantitative values of sequence reads for a subregion may be part of intensive sequence analysis. Providing quantitative values of sequence reads for a subregion may be part of intensive sequence analysis of nucleic acids captured by probe oligonucleotides. In some embodiments, classification of the presence or absence of copy number variation in a subchromosomal region is provided according to the presence or absence of copy number variation segments. In some embodiments, classification of the presence or absence of copy number variation in a subchromosomal region is provided according to quantitative values of sequence reads for subregions within the subchromosomal region. In some embodiments, classification of the presence or absence of copy number variation in a subchromosomal region is provided according to the presence or absence of copy number variation segments and according to quantitative values of sequence reads for subregions within the subchromosomal region.
[0026] In some embodiments, the classification of the presence or absence of copy number variation in a subchromosomal region includes providing a quantitative value of sequence reads for a subregion within the subchromosomal region, the subregion comprising a predetermined set of genomic subregions. A predetermined set of genomic subregions may be identified according to one or more precision measures for multiple samples (e.g., multiple samples in a training set). Generally, each sample in the set of multiple samples (e.g., a training set) is classified as having copy number variation in the subchromosomal region of interest. Samples in the set of multiple samples may be obtained from one or more subjects known to have copy number variation, and / or generated by adding genomic DNA with copy number variation to a reference sample, and / or generated according to in silico modeling. Having copy number variation in a subchromosomal region of interest may include copy number variation identified in genomic coordinates within the subchromosomal region of interest, copy number variation identified in genomic coordinates overlapping the subchromosomal region of interest, copy number variation identified in genomic coordinates adjacent to the subchromosomal region of interest (e.g., within approximately 1 megabase of the subchromosomal region of interest), and so on. Copy number variations in a set of multiple samples may include duplications, microduplications, deletions, and microdeletions. Duplications and deletions can be of any size, but microduplications and microdeletions typically refer to duplications and deletions smaller than 5 million bases that are generally too small to be detected by conventional cytogenetic methods or high-resolution karyotype analysis.
[0027] A precision measure for multiple samples may include any suitable precision measure for determining the presence or absence of copy number variation for multiple samples. Precision measures may include sensitivity, specificity, standard deviation, median absolute deviation (MAD), determinism, confidence, uncertainty, coefficient of variation (CV), confidence level, confidence interval (e.g., approximately 95% confidence interval), standard score (e.g., z score), chi-value, phi-value, t-test result, p-value, ploidy value, fitted minority proportion, area ratio, median level, etc., or combinations thereof. In some embodiments, the precision measure includes sensitivity.
[0028] Typically, each of the above-mentioned multiple samples (e.g., multiple samples in a training set) has known copy number variations, so the accuracy of detecting copy number variations can be evaluated. In some embodiments, the accuracy of detecting copy number variations for the multiple samples can be optimized. In some embodiments, the accuracy of detecting copy number variations for the multiple samples can be optimized by identifying a set of genomic subsets that provides the optimal accuracy measure for classifying the presence or absence of copy number variations for the multiple samples. Where disclosed herein, the term “optimal accuracy” means an accuracy measure that is equal to or higher than a predetermined (predetermine) threshold, which is considered the minimum requirement for detecting the presence or absence of copy number variations with reasonable accuracy. Those skilled in the art can readily determine what the predetermined threshold is for any particular accuracy measure required for a particular assay. In some embodiments, the accuracy of detecting copy number variations for multiple samples can be optimized by identifying a set of genomic subsets that provides the optimal sensitivity for classifying the presence or absence of copy number variations for the multiple samples. In some embodiments, the set of genomic subsets that provides the optimal accuracy measure (e.g., optimal sensitivity) is referred to as a predetermined genomic subset or a predetermined subregion. In some embodiments, a set of genomic subregions that provides an optimal precision measure (e.g., optimal sensitivity) is identified by a process that includes: 1) providing a plurality of candidate subregions within a subchromosome region of interest (e.g., a subchromosome region with possible copy number variation); 2) providing one or more precision measures (e.g., sensitivity values) for each of the plurality of candidate subregions for a plurality of samples (e.g., in a training set); and 3) identifying a set of genomic subregions in which the subregion provides an optimal precision measure (e.g., optimal sensitivity) according to that one or more precision measure. The plurality of candidate subregions provided to identify a set of genomic subregions that provides an optimal precision measure typically include subregions having one or more distinct genomic coordinates.For example, each candidate subregion may have a unique genomic coordinate at its 5' end, or a unique genomic coordinate at its 3' end, or a unique genomic coordinate at both its 5' and 3' ends. The candidate subregions may be the same length as each other, or of different lengths, or a combination of both.
[0029] In some embodiments, one or more precision measures include a sensitivity measure. Sensitivity may be determined as the number or percentage of samples identified as having copy number variation, where the samples come from a set of multiple samples having copy number variation. In some embodiments, the sensitivity for classifying each of the multiple samples (e.g., in a training set) as having copy number variation in the subchromosomal region of interest is at least about 70%. For example, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region may be at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 75%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 80%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 85%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 90%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 95%. In some embodiments, the sensitivity for classifying each of several samples (e.g., in a training set) as having a copy number variation in the target subchromosome region is at least about 97%.
[0030] In some embodiments, the classification of the presence or absence of copy number variation in a subchromosomal region includes providing a quantitative value of sequence reads for the subregion (e.g., the subregions described above). The quantitative value of sequence reads for a subregion may be a sequence read count (e.g., direct sum of read counts, raw read count, normalized read count, filtered read count, read density, weighted read count, read count ratio, mean read count, mean read count, adjusted read count, etc., and combinations thereof). In some embodiments, the quantitative value of sequence reads for a subregion is a normalized quantitative value of sequence reads generated by a normalization process. The normalization process may include any preferred normalization that normalizes GC bias and / or other biases. Examples of certain normalization processes are described herein. In some embodiments, the normalization process includes LOESS normalization. In some embodiments, the normalization process includes principal component normalization. The classification of the presence or absence of copy number variation in a subchromosomal region may be based on the change in the quantitative value of sequence reads relative to a reference sample set. For the purposes of this disclosure, the reference sample set may be any sample identified as not having the copy number variation to be detected in the test sample. The reference sample may originate from a similar tissue type and / or similar population type of subject that does not have the copy number variation.
[0031] In some embodiments, the quantification of sequence reads for a subregion is a standard score. In some embodiments, the quantification of sequence reads for a subregion is a z score. The z score may be for a subregion, or it may be assigned to each genomic portion contained within the subregion. The z scores are as follows: Z SUB =(SUB scq -SUB mcq ) / MAD It can be generated for the subregion according to (Z SUB ).
[0032] In the formula, SUBscq is the test sample count quantification value of the sub-region (e.g., SUB scq can be the result of dividing the normalized total count in the sub-region for the test sample by the normalized total count of the autosome); SUB mcq is the median of the count quantification values for the sub-region generated for the reference sample set; MAD is the median absolute deviation determined for the count quantification values of the sub-region for the reference sample set. In a particular case, SUB mcq is the mean of the count quantification values for the sub-region generated for the reference sample set; the denominator of the above equation is the standard deviation determined for the count quantification values of the sub-region for the reference sample set. In a particular case, SUB scq can be the result of dividing the total count in the sub-region for the test sample by the total count of the autosome. The total count of the autosome can be normalized (e.g., can be GC-normalized), filtered (e.g., repeat regions can be filtered out, low mapping regions can be filtered out, and / or other regions can be filtered out as described herein), or normalized and filtered. In a particular case, SUB scqThis may be the result of dividing the total count in a subregion (e.g., normalized total count) by the total count for a genome subset of the test sample (e.g., normalized total count). Genomic subsets may include, for example, all autosomes, parts of all autosomes, certain autosomes, parts of certain autosomes, and combinations thereof. The reference sample set may include samples classified as having no copy number variation. In some embodiments, the reference sample consists of samples classified as having no copy number variation. Therefore, in some embodiments, the reference sample includes or consists of samples in which each chromosome and each chromosomal region being tested is euploid. The reference sample may be derived from human subjects. In some embodiments, the reference sample is derived from female subjects. In some embodiments, the reference sample is derived from male subjects. In some embodiments, the reference sample is derived from both male and female subjects. The reference sample may include samples from one subject or samples from multiple subjects. A reference sample may contain one reference sample, but often it contains multiple samples. For example, a reference sample may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more samples. Other quantitative values may be used instead of z-scores; non-restrictive examples include normal scores, z-values, standardized variables, and t-statistics.
[0033] In some embodiments, the presence or absence of copy number variation in a subregion is classified according to a z-score cutoff. The z-score cutoff may be determined according to a preferred level of sensitivity and / or specificity for determining the presence or absence of copy number variation in a test sample. In some embodiments, the z-score cutoff value is set to an absolute value of about 2 to about 4. For example, the z-score cutoff value may be set to an absolute value of about 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0. In some embodiments, the z-score cutoff value is set to an absolute value of about 3 to about 5. For example, the z-score cutoff value can be set to an absolute value of approximately 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0. In some embodiments, the z-score cutoff value is set to an absolute value of approximately 3.9 to approximately 4.0. For example, the z-score cutoff value can be set to an absolute value of approximately 3.90, 3.91, 3.92, 3.93, 3.94, 3.95, 3.96, 3.97, 3.98, 3.99, or 4.0. In some embodiments, the z-score cutoff value is set to an absolute value of approximately 3.95. The presence or absence of copy number variation in a test sample can be determined if the absolute value of one or more z-scores for a given subregion is greater than a selected cutoff value. In some embodiments, the classification of the presence or absence of copy number variation in a subchromosomal region for a test sample is provided according to the quantitative value (e.g., z-score) of the sequence reads for the subregion. In some embodiments, the classification of the presence of a deletion in a subchromosomal region is determined if the z-score generated using the method described herein is less than -3, less than -3.2, or less than -3.5, for example, less than -3.95. In some embodiments, the classification of the presence of a duplication in a subchromosomal region is made if the z-score is greater than 3, greater than 3.2, or greater than 3.5, for example, greater than 3.95.
[0034] In some embodiments, the classification of the presence or absence of copy number variation in a subchromosomal region includes identifying the presence or absence of a copy number variation segment. In some embodiments, the copy number variation segment is identified using a method that includes a segmentation process. A method that includes a segmentation process may include a decision analysis, such as a decision analysis as described herein. For example, a decision analysis may include applying one or more methods that result in one or more outcomes, evaluations of the outcomes and a set of decisions, based on the possible consequences of those outcomes, evaluations and / or those decisions, and terminating at some critical stage of the process in which a final decision is made. In some embodiments, the decision analysis is a decision tree. In some embodiments, the presence or absence of a copy number variation segment is identified according to a segmentation process or a decision analysis that includes a segmentation process.
[0035] In some embodiments, a segmentation process is applied to identify segments (e.g., segments spanning copy number variations; copy number variation segments). Any suitable segmentation process may be used, including, but is not limited to, the circular binary segmentation (CBS) process. CBS generally works by repeatedly dividing a single chromosome into regions of equal copy number using likelihood ratio statistics. CBS is described, for example, in Olshen et al. (2004) Biostatistics 5:557-72; Venkatraman et al. (2007) Bioinformatics 23:657-63; Lai et al. (2005) Bioinformatics 21:3763-70; and Willenbrock et al. (2005) Bioinformatics 21:4084-91. Other processes can be used instead of or in addition to CBS, and non-limiting examples include wavelet segmentation (e.g., Haar wavelet segmentation), Fourier transform, sliding window z-scores, and Markov chain models.
[0036] In some embodiments, the method for classifying the presence or absence of copy number variations involves genome-wide analysis, i.e., analysis based on circular binary segmentation (CBS) to find events within a genomic window encompassing a target region, e.g., 22q11.2, e.g., edges of microdeletions or microduplications. CBS is useful for detecting small deletions. In some embodiments, the method for classifying the presence or absence of copy number variations involves intensive analysis, i.e., analysis using a predetermined region within the target region. Generally, when the test sample contains a low fetal proportion, e.g., less than 10%, intensive sequencing analysis is more reliable and / or more sensitive, while when the test sample contains a high fetal proportion, e.g., greater than 10%, genome-wide sequencing analysis may be more sensitive and therefore preferred. An illustrative embodiment is shown in Figure 2. In certain embodiments, the above method utilizes both genome-wide analysis and intensive sequencing analysis, maximizing sensitivity by employing the edge detection capabilities of CBS, thereby enabling the identification of small deletions and improving sensitivity at low fetal proportions through intensive sequencing analysis.
[0037] In some embodiments, quantitative values are generated for copy number variant segments identified by the segmentation process. In some embodiments, the segmentation process generates quantitative values for copy number variant segments. The quantitative values for copy number variant segments may include quantitative values for sequence reads. The quantitative values for sequence reads for copy number variant segments may be sequence read counts (e.g., direct sum of read counts, raw read counts, normalized read counts, filtered read counts, read density, weighted read counts, read count ratios, average read counts, average read counts, adjusted read counts, etc., and combinations thereof). In some embodiments, the quantitative values for sequence reads for copy number variant segments are normalized sequence read quantitative values generated by the normalization process. The normalization process may include any preferred normalization that normalizes GC bias and / or other biases. Examples of certain normalization processes are described herein. In some embodiments, the normalization process includes LOESS normalization. In some embodiments, the normalization process includes principal component normalization.
[0038] In some embodiments, the quantitative value for a copy number variant segment is a standard score. In some embodiments, the quantitative value for a copy number variant segment is a z-score. The z-score may be for a segment and may be assigned to each genomic portion contained within the segment. The z-score is as follows: Z SEG =(SEG scq -SEG mcq ) / MAD This can be generated for copy number variant segments according to (Z SEG ).
[0039] In the formula, SEG scq This is the quantitative value of the test sample count for a segment (e.g., SEG scqThis may be the result of dividing the normalized total count in the segment for the test sample by the normalized total count of the autosomes); SEG mcq SEG is the median of the count quantifiable values for the segment generated relative to the reference sample set; MAD is the median absolute deviation determined relative to the count quantifiable values of the segment relative to the reference sample set. In a particular case, SEG mcq is the mean of the count quantified values for the segments generated relative to the reference sample set; the denominator of the above equation is the standard deviation determined for the count quantified values of the segments relative to the reference sample set. In a particular case, SEG scq This may be the result of dividing the total count in a subregion of the test sample by the total count of the autosomes. The total count of the autosomes may be normalized (e.g., GC normalized), filtered (e.g., repeat regions may be filtered and excluded, low-mapping regions may be filtered and excluded, and / or other regions may be filtered and excluded as described herein), or normalized and filtered. In certain cases, SEG scq This may be the result of dividing the total count in a subregion of the test sample (e.g., normalized total count) by the total count for a genome subset (e.g., normalized total count). Examples of genome subsets include all autosomes, parts of all autosomes, certain autosomes, parts of certain autosomes, and combinations thereof. The reference sample set may be any suitable reference set and may include the reference sample sets described herein.
[0040] Non-limiting examples of methodologies useful for generating z-score copy number quantitative values based on segmentation (e.g., CBS) are described in Zhao et al., Clin. Chem. 61:4:608-616 (2015); Lefkowitz et al., American Journal of Obstetrics & Gynecology 1.e1 (2016); and International Patent Application No. PCT / US2014 / 039389 (filed May 23, 2014, and published November 27, 2014 as WO2014 / 190286). Other normalized CNV quantitative values may be used instead of z-scores, non-limiting examples of which include normalized scores, z-values, standardized variables, and t-statistics.
[0041] In some embodiments, the presence or absence of copy number variation in a segment is classified according to a z-score cutoff. The z-score cutoff may be determined according to a preferred level of sensitivity and / or specificity for determining the presence or absence of copy number variation in a test sample. In some embodiments, the z-score cutoff value is set to an absolute value of about 2 to about 4. For example, the z-score cutoff value may be set to an absolute value of about 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0. In some embodiments, the z-score cutoff value is set to an absolute value of about 3 to about 5. For example, the z-score cutoff value can be set to an absolute value of approximately 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0. In some embodiments, the z-score cutoff value is set to an absolute value of approximately 3.9 to approximately 4.0. For example, the z-score cutoff value can be set to an absolute value of approximately 3.90, 3.91, 3.92, 3.93, 3.94, 3.95, 3.96, 3.97, 3.98, 3.99, or 4.0. In some embodiments, the z-score cutoff value is set to an absolute value of approximately 3.95. The presence or absence of copy number variation in a test sample can be determined if the absolute value of one or more z-scores for a given segment is greater than a selected cutoff value. In some embodiments, the classification of the presence or absence of copy number variation in a subchromosomal region of a test sample is provided according to quantitative values (e.g., z-scores) for the copy number variation segment.
[0042] In some embodiments, the classification of the presence or absence of copy number mutations in a subchromosomal region of a test sample is provided according to quantitative values (e.g., z-scores) for copy number mutation segments and quantitative values (e.g., z-scores) for sequence reads to subregions. In some embodiments, the classification of the presence or absence of copy number mutations in a subchromosomal region of a test sample is provided according to quantitative values (e.g., z-scores) for copy number mutation segments or quantitative values (e.g., z-scores) for sequence reads to subregions. Thus, in certain cases, the classification is provided according to quantitative values (e.g., z-scores) for both segments and subregions, and in certain cases, the classification is provided according to either quantitative values (e.g., z-scores) for segments or quantitative values (e.g., z-scores) for subregions.
[0043] In some embodiments, a segment includes a first genome subset, and a subregion includes a second genome subset. In some embodiments, the first and second genome subsets include the same genome portion. In some embodiments, the first and second genome subsets consist of the same genome portion. In some embodiments, the first and second genome subsets include different genome portions. In some embodiments, the first and second genome subsets include some identical genome portions and some different genome portions. In some embodiments, the second genome subset is a subset of the first genome subset. In some embodiments, the first genome subset is a subset of the second genome subset. In some embodiments, the second genome subset overlaps with the first genome subset. In some embodiments, the second genome subset partially overlaps with the first genome subset. In some embodiments, the second genome subset includes fewer genome portions than the first genome subset. In some embodiments, the second set of genome segments contains more genome segments than the first set of genome segments.
[0044] In some embodiments, the methods described herein include a step of classifying the presence or absence of microduplications in subchromosomal regions. Microduplications may be duplications in chromosomes selected from chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X chromosome, and Y chromosome. In some embodiments, the methods described herein include a step of classifying the presence or absence of microdeletions in subchromosomal regions. Microdeletions may be deletions in chromosomes selected from chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X chromosome, and Y chromosome. In some embodiments, a microdeletion is a deletion in a genomic region or a portion of a genomic region selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23. In some embodiments, a microdeletion or microduplication is associated with a disease or syndrome. Examples of syndromes that may be associated with a particular microdeletion and / or microduplication include 1p36 syndrome, DiGeorge syndrome, Prader-Willi syndrome, Angelman syndrome, Langer-Giedion syndrome, Jacobsen syndrome, Wolf-Hirschhorn syndrome, Miller-Dieker syndrome, and Williams-Buren syndrome. A non-limiting list of known and / or possible associations between copy number variations in a particular genomic region and syndromes is provided in Table 1 below. [Table 1]
[0045] In some embodiments, copy number variations in subchromosomal regions are characterized by their size (i.e., length). The length of a copy number variation in a subchromosomal region refers to the number of consecutive nucleotide bases that are deleted (e.g., in the case of a microdeletion) or duplicated (e.g., in the case of a microduplication). In some embodiments, the length of a copy number variation in a subchromosomal region is about 1 megabase or less. For example, the length of a copy number variation in a subchromosomal region may be about 900 kilobases (kb), 800kb, 700kb, 600kb, 500kb, 400kb, 300kb, 200kb, or 100kb. In some embodiments, the length of a copy number variation in a subchromosomal region is about 1 megabase to about 40 megabases. For example, the length of copy number variations in subchromosomal regions can be approximately 1 megabase to 2 megabases, 1 megabase to 3 megabases, 1 megabase to 4 megabases, 1 megabase to 5 megabases, 1 megabase to 6 megabases, 1 megabase to 7 megabases, 1 megabase to 8 megabases, 1 megabase to 9 megabases, 1 megabase to 10 megabases, 1 megabase to 11 megabases, 1 megabase to 12 megabases, 1 megabase to 13 megabases, 1 megabase to 14 megabases, 1 megabase to 15 megabases, 1 megabase to 16 megabases, 1 megabase to 17 megabases, 1 megabase to 18 megabases, 1 megabase to 19 megabases, 1 megabase to 20 megabases, 1 megabase to 25 megabases, 1 megabase to 30 megabases, 1 megabase to 35 megabases, or 1 megabase to 40 megabases. In some embodiments, the length of the copy number variation in the subchromosomal region is approximately 1 megabase to approximately 20 megabases. In some embodiments, the length of the copy number variation in the subchromosomal region is approximately 1 megabase to approximately 10 megabases. In some embodiments, the length of the copy number variation in the subchromosomal region is approximately 1 megabase to approximately 7 megabases.
[0046] In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample can be classified with a sensitivity of at least about 70%. For example, the presence or absence of copy number variations in subchromosomal regions of a test sample can be classified with a sensitivity of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample can be classified with a sensitivity of at least about 75%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample can be classified with a sensitivity of at least about 80%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample is classified with a sensitivity of at least about 85%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample is classified with a sensitivity of at least about 90%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample is classified with a sensitivity of at least about 95%. In some embodiments, the presence or absence of copy number variations in subchromosomal regions of a test sample is classified with a sensitivity of at least about 97%.
[0047] In some embodiments, the presence or absence of copy number variations in subchromosomal regions to a test sample is classified with at least about 90% specificity. For example, the presence or absence of copy number variations in subchromosomal regions to a test sample can be classified with at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% specificity. In some embodiments, the presence or absence of copy number variations in subchromosomal regions to a test sample is classified with at least about 99% specificity. In some embodiments, the presence or absence of copy number variations in subchromosomal regions to a test sample is classified with at least about 99.9% specificity. In some embodiments, the presence or absence of copy number variations in subchromosomal regions in a test sample is classified with approximately 100% specificity.
[0048] In some embodiments, the nucleic acids in the test sample are derived from the test subject. In some embodiments, the nucleic acids in the test sample include circulating cell-free nucleic acids. In some embodiments, the circulating cell-free nucleic acids are derived from the plasma or serum of the test subject. In some embodiments, the test subject is male. In some embodiments, the test subject is a human male. In some embodiments, the test subject is female. In some embodiments, the test subject is a human female. In some embodiments, the test subject is pregnant. In some embodiments, the nucleic acids in the test sample include maternal nucleic acids and fetal nucleic acids. In some embodiments, the proportion of fetal nucleic acids in the test sample is less than about 25%. For example, the proportion of fetal nucleic acids in a test sample may be approximately 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. In some embodiments, the proportion of fetal nucleic acids in a test sample is less than approximately 10%. In some embodiments, the proportion of fetal nucleic acids in a test sample is less than approximately 5%. In some embodiments, the test subject is a cancer patient or a subject being tested or screened for cancer. In some embodiments, the nucleic acids in the test sample include patient / host nucleic acids and tumor nucleic acids or nucleic acids derived from cancer cells. In some embodiments, the proportion of tumor / cancer nucleic acids in a test sample is less than approximately 25%. For example, the proportion of tumor / cancer nucleic acids in a test sample may be approximately 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. In some embodiments, the proportion of tumor / cancer nucleic acids in a test sample is less than approximately 10%. In some embodiments, the proportion of tumor / cancer nucleic acids in a test sample is less than approximately 5%.
[0049] sample Systems, methods, and products for analyzing nucleic acids are provided herein. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. Nucleic acid fragments may be referred to as nucleic acid templates, and these terms may be used interchangeably herein. A mixture of nucleic acids may comprise two or more nucleic acid fragment species having the same or different nucleotide sequences, different fragment lengths, different origins (e.g., genomic origin, fetal origin vs. maternal origin, cellular or tissue origin, cancer origin vs. non-cancerous origin, tumor origin vs. non-tumor origin, sample origin, subject origin, etc.) or combinations thereof.
[0050] Nucleic acids or nucleic acid mixtures used in the systems, methods, and products described herein are often isolated from samples obtained from a subject (e.g., a test subject). The subject can be any living or non-living organism, including, but not limited to, humans, non-human animals, plants, bacteria, fungi, protists, or pathogens. Any human or non-human animal can be selected, including, for example, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle (e.g., cows), horses (e.g., horses), goats and sheep (e.g., sheep, goats), pigs (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursids (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject may be male or female (e.g., female, pregnant). The subject may be of any age (e.g., embryo, fetus, infant, child, adult). The subject may be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. In some embodiments, the test subject is female. In some embodiments, the test subject is human female. In some embodiments, the test subject is male. In some embodiments, the test subject is human male.
[0051] Nucleic acids can be isolated from any type of suitable biological specimen or sample (e.g., a test sample). A sample or test sample may be any specimen isolated from or obtained from a subject or a part thereof (e.g., a human subject, a pregnant woman, a cancer patient, a fetus, a tumor). Non-limiting examples of specimens include, but are not limited to, fluids or tissues derived from the subject, including, blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar lavage fluid, gastric lavage fluid, peritoneal lavage fluid, tubal lavage fluid, ear lavage fluid, arthroscopic lavage fluid), biopsy samples (e.g., preimplantation embryos; cancer biopsy materials), celocentesis samples, cells (blood cells, placental cells, embryonic cells, or fetal cells, fetal nucleated cells or fetal cellular remnants, normal cells, abnormal cells (e.g., cancer cells)) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), female reproductive tract lavage fluid, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, milk, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab derived from the subject. The fluid or tissue sample from which nucleic acids are extracted may be cell-free (e.g., cell-free). In some embodiments, the fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancer cells may be present in the sample.
[0052] The sample may be a liquid sample. A liquid sample may contain extracellular nucleic acids (e.g., circulating cell-free DNA). Non-limiting examples of liquid samples include blood or blood products (e.g., serum, plasma, etc.), urine, biopsy samples (e.g., liquid biopsy material for detecting cancer), the liquid samples described above, or combinations thereof. In certain embodiments, the sample is a liquid biopsy material, which broadly refers to the evaluation of a liquid sample derived from a subject for the presence or absence, progression or remission of a disease (e.g., cancer). Liquid biopsy material may be used together with or as a substitute for solid biopsy material (e.g., tumor biopsy material). In certain cases, extracellular nucleic acids are analyzed in the liquid biopsy material.
[0053] In some embodiments, the biological sample may be blood, plasma, or serum. The term “blood” encompasses whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc., as conventionally defined. Blood or its fractions often contain nucleosomes. Nucleosomes contain nucleic acids and may be cell-free or intracellular. Blood also contains buffy coat. Buffy coat may be isolated by using a Ficol gradient. Buffy coat may contain leukocytes (e.g., white blood cells, T cells, B cells, platelets, etc.). Plasma refers to the fraction of whole blood resulting from centrifugation of blood treated with an anticoagulant. Serum refers to the watery portion of the fluid remaining after a blood sample has clotted. Fluid or tissue samples are often collected according to standard protocols commonly followed by hospitals or clinics. In the case of blood, an appropriate amount of peripheral blood (e.g., 3-40 ml, 5-50 ml) is often collected, which may be stored according to standard procedures before or after preparation.
[0054] Analysis of nucleic acids found in the blood of a subject may be performed, for example, using whole blood, serum, or plasma. Analysis of fetal DNA found in maternal blood may be performed, for example, using whole blood, serum, or plasma. Analysis of tumor DNA found in patient blood may be performed, for example, using whole blood, serum, or plasma. Methods for preparing serum or plasma from blood obtained from a subject (e.g., maternal subject; cancer patient) are known. For example, the blood of a subject (e.g., blood of a pregnant woman; blood of a cancer patient) may be placed in a tube containing EDTA or a commercially available specialized device such as a Vacutainer SST (Becton Dickinson, Franklin Lakes, NJ) to prevent blood clotting, and then plasma can be obtained from the whole blood by centrifugation. Serum may be obtained with or without blood clotting after centrifugation. When centrifugation is used, it is usually performed at an appropriate speed, for example, 1,500 to 3,000 × g, but is not limited to this. Plasma or serum may be subjected to further centrifugation steps and then transferred to a new tube for nucleic acid extraction. In addition to the cell-free portion of whole blood, nucleic acids can also be recovered from the cell fraction concentrated in the buffy coat portion obtained after centrifugation and plasma removal of a whole blood sample derived from the subject.
[0055] Samples may be heterogeneous. For example, a sample may contain more than one cell type and / or one or more nucleic acid species. In some cases, a sample may contain (i) fetal and maternal cells, (ii) cancer and non-cancerous cells, and / or (iii) pathogenic and host cells. In some cases, a sample may contain (i) cancer and non-cancerous nucleic acids, (ii) pathogen and host nucleic acids, (iii) fetal and maternal nucleic acids, and / or more generally, (iv) mutant and wild-type nucleic acids. In some cases, a sample may contain minority nucleic acid species and majority nucleic acid species, as described in more detail below. In some cases, a sample may contain cells and / or nucleic acids from a single subject or from multiple subjects.
[0056] cell type As used herein, “cell type” refers to a type of cell that can be distinguished from other types of cells. Extracellular nucleic acids may include nucleic acids derived from several different cell types. Non-limiting examples of cell types that can lead nucleic acids to circulating cell-free nucleic acids include liver cells (e.g., hepatocytes), lung cells, spleen cells, pancreatic cells, colon cells, skin cells, bladder cells, eye cells, brain cells, esophageal cells, head cells, neck cells, ovarian cells, testicular cells, prostate cells, placental cells, epithelial cells, endothelial cells, adipocytes, kidney / renal cells, cardiac cells, muscle cells, blood cells (e.g., leukocytes), central nervous system (CNS) cells, and combinations of the aforementioned cells. In some embodiments, cell types that lead nucleic acids to circulating cell-free nucleic acids being analyzed include leukocytes, endothelial cells, and hepatocytes (liver cells). As will be described in more detail herein, various cell types may be screened as part of identifying and selecting nucleic acid loci whose marker status is the same or substantially the same as that of cell types in subjects with medical symptoms and cell types in subjects without medical symptoms.
[0057] Certain cell types may remain the same or substantially the same in subjects with and without medical symptoms. In non-limiting cases, the number of living or viable cells of a particular cell type may decrease in certain cytopathic symptoms, and living, viable cells may not be altered or significantly altered in subjects with that medical symptom.
[0058] Certain cell types may be modified as part of a medical condition, sometimes possessing one or more characteristics that differ from their original state. In non-limiting examples, certain cell types may, as part of a cancerous condition, grow at a rate faster than normal, become cancerous in cells with different morphologies, become cancerous in cells expressing one or more different cell surface markers, and / or become part of a tumor. In embodiments where a certain cell type (i.e., progenitor cells) is modified as part of a medical condition, the marker status for each of the one or more markers being assayed is often the same or substantially the same for that particular cell type in subjects with the medical condition and for that particular cell type in subjects without the medical condition. Thus, the term “cell type” may sometimes refer to a type of cell in a subject without a medical condition and a modified version of that cell in a subject with the medical condition. In some embodiments, “cell type” refers only to the progenitor cell and not to a modified version arising from the progenitor cell. “Cell type” may sometimes refer to the progenitor cell and the modified cell arising from the progenitor cell. In such embodiments, the state of the marker being analyzed is often the same or substantially the same as the cell type in a subject with a certain medical condition and the cell type in a subject without that medical condition.
[0059] In certain embodiments, the cell type is cancer cells. Certain types of cancer cells include, for example, leukemia cells (e.g., acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphoblastic leukemia); cancerous kidney / renal cells (e.g., renal cell carcinoma (clear cell, type 1 papillary, type 2 papillary, chromophobe, ampullae, collecting duct), renal adenocarcinoma, adrenal tumor, Wilms' tumor, transitional cell carcinoma); brain tumor cells (e.g., acoustic neuroma, astrocytoma (grade I: pilocytic astrocytoma, grade II)). Examples include low-grade astrocytoma, grade III undifferentiated astrocytoma, grade IV glioblastoma (GBM), chordoma, CNS lymphoma, craniopharyngioma, glioma (brainstem glioma, ependymoma, mixed glioma, optic glioma, subependymoma), medulloblastoma, meningioma, metastatic brain tumor, oligodendroglioma, pituitary tumor, primitive neuroectodermal tumor (PNET), schwannoma, juvenile pilocytic astrocytoma (JPA), pineal tumor, and rhabdoid tumor.
[0060] Different cell types may be distinguished by any preferred features, including, but not limited to, one or more different cell surface markers, one or more different morphological features, one or more different functions, one or more different protein (e.g., histone) modifications, and one or more different nucleic acid markers. Non-limiting examples of nucleic acid markers include single nucleotide polymorphisms (SNPs), methylation status of nucleic acid loci, short tandem repeats, insertions (e.g., microinsertions), deletions (microdeletions), and combinations thereof. Non-limiting examples of protein (e.g., histone) modifications include acetylation, methylation, ubiquitination, phosphorylation, SUMOylation, and combinations thereof.
[0061] As used herein, the term “related cell type” refers to a cell type that shares several characteristics with another cell type. In related cell types, 75% or more of the cell surface markers may be common to that cell type (for example, approximately 80%, 85%, 90%, or 95% or more of the cell surface markers may be common to the related cell type).
[0062] Nucleic acid Methods for analyzing nucleic acids are provided herein. The terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid fragment,” and “nucleic acid template” may be used interchangeably throughout this disclosure. These terms refer to nucleic acids of any composition from, for example, DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., message RNA (mRNA), small inhibitory RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA analogs or RNA analogs (e.g., those containing base analogs, sugar analogs, and / or non-natural backbones), RNA / DNA hybrids, and polyamide nucleic acids (PNA), all of which may be in single-stranded or double-stranded form and may include known analogs of natural nucleotides that, unless otherwise limited, may function in a manner similar to naturally occurring nucleotides. Nucleic acids may be, or may be derived from, plasmids, phages, viruses, bacteria, autonomous replication sequences (ARS), mitochondria, centromeres, artificial chromosomes, chromosomes, or other nucleic acids that can or may be replicated in vitro or in host cells, cells, cell nuclei, or the cytoplasm of cells. In some embodiments, the template nucleic acid may be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless specifically limited, this term encompasses nucleic acids including known analogs of native nucleotides that have similar binding properties to a reference nucleic acid and are metabolized in a similar manner to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence implicitly includes its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologues, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as explicitly indicated sequences. Specifically, degenerate codon substitution can be achieved by creating a sequence in which the third position of one or more selected (or all) codons is replaced with a mixed base and / or deoxyinosine residue.The term nucleic acid is used interchangeably with gene locus, gene, cDNA, and mRNA encoded by a gene. The term may also include single-stranded polynucleotides ("sense" or "antisense," "plus" or "minus" strands, "forward" or "reverse" reading frames) and double-stranded polynucleotides as equivalents, derivatives, variants, and analogs of RNA or DNA synthesized from nucleotide analogs. The term "gene" refers to a region of DNA involved in the formation of a polypeptide chain; the term generally includes the regions before and after the coding region (leaders and trailers) involved in the transcription / translation and regulation of the gene product, as well as the intervening sequences (introns) between individual coding regions (exons). A nucleotide or base generally refers to the purine and pyrimidine molecular units of a nucleic acid (e.g., adenine (A), thymine (T), guanine (G), and cytosine (C)). In the case of RNA, the base thymine is replaced by uracil. The length or size of a nucleic acid can be expressed as the number of bases.
[0063] Nucleic acids can be single-stranded or double-stranded. For example, single-stranded DNA can be produced by denaturing double-stranded DNA, for example, by heating or alkaline treatment. In certain embodiments, the nucleic acid is a D-loop structure formed by strand insertion of an oligonucleotide or DNA-like molecule, such as peptide nucleic acid (PNA), into a double-stranded DNA molecule. D-loop formation can be promoted using methods known in the art, for example, by adding E. coli RecA protein and / or changing the salt concentration.
[0064] The nucleic acids provided for the processes described herein may include nucleic acids derived from one or more samples (for example, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more samples).
[0065] Nucleic acids can be obtained from one or more sources (e.g., biological samples, blood, cells, serum, plasma, buffy coat, urine, lymph, skin, soil, etc.) by methods known in the art. Any suitable method can be used to isolate, extract, and / or purify DNA from a biological sample (e.g., blood or blood products), non-limiting examples of which include methods for DNA preparation (e.g., those described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3rd ed., 2001), various commercially available reagents or kits, e.g., Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep TM Blood DNA Isolation Kit (Promega, Madison, Wis.) and GFX TM Examples include Genomic Blood DNA Purification Kits (Amersham, Piscataway, NJ) or combinations thereof.
[0066] In some embodiments, nucleic acids are extracted from cells using cell lysis procedures. Cell lysis procedures and reagents are known in the art and can generally be carried out by chemical lysis methods (e.g., washing agents, hypotonic solutions, enzymatic procedures, etc., or combinations thereof), physical lysis methods (e.g., French press, sonication, etc.), or electrolysis. Any suitable lysis procedure can be used. For example, chemical methods generally involve using a lysis agent to destroy the cells, extracting nucleic acids from those cells, and then treating them with chaotropic salts. Physical methods such as grinding after freezing / thawing and the use of cell presses are also useful. In some cases, high-salt lysis procedures and / or alkaline lysis procedures may be used.
[0067] In certain embodiments, nucleic acids may include extracellular nucleic acids. The term “extracellular nucleic acid,” as used herein, may refer to nucleic acids isolated from substantially cell-free sources, and may also be referred to as “cell-free” nucleic acids, “circulating cell-free nucleic acids” (e.g., CCF fragments, ccfDNA), and / or “cell-free circulating nucleic acids.” Extracellular nucleic acids may be present in blood (e.g., the blood of a human subject) and can be obtained from that blood. Extracellular nucleic acids often do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources for extracellular nucleic acids include blood, plasma, serum, and urine. As used herein, the term “obtaining a cell-free circulating sample nucleic acid” includes obtaining a sample directly (e.g., recovering a sample, e.g., a test sample) or obtaining a sample from another person who has recovered a sample. While not limited to theory, extracellular nucleic acids may be products of cellular apoptosis and cellular destruction, which often form the basis of a range of extracellular nucleic acids (e.g., a “ladder”). In some embodiments, the sample nucleic acid derived from the test subject is a circulating cell-free nucleic acid. In some embodiments, the circulating cell-free nucleic acid is derived from the test subject's plasma or serum.
[0068] Extracellular nucleic acids may contain various types of nucleic acids and are therefore referred to herein as "heterogeneous" in certain embodiments. For example, serum or plasma from a person with cancer may contain nucleic acids derived from cancer cells (e.g., tumors, neoplasias) and nucleic acids derived from non-cancerous cells. In another example, serum or plasma from a pregnant woman may contain maternal nucleic acids and fetal nucleic acids. In some cases, cancer nucleic acids or fetal nucleic acids may account for approximately 5% to 50% of the total nucleic acids (for example, approximately 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acids are cancer nucleic acids or fetal nucleic acids).
[0069] At least two different nucleic acid species may exist as extracellular nucleic acids in different amounts, and these are sometimes referred to as minority species and majority species. In certain cases, minority nucleic acids originate from affected cell types (e.g., cancer cells, wasting cells, cells attacked by the immune system). In certain embodiments, gene mutations or genetic alterations (e.g., copy number changes, copy number variations, single nucleotide changes, single nucleotide variations, chromosomal changes and / or translocations) are determined for minority nucleic acids. In certain embodiments, gene mutations or genetic alterations are determined for majority nucleic acids. In general, the terms “minority” and “majority” are not intended to be strictly defined in any respect. In one embodiment, nucleic acids considered “minority” may, for example, be present in amounts ranging from at least about 0.1% to less than 50% of the total nucleic acids in the sample. In some embodiments, minority nucleic acids may be present in amounts ranging from at least about 1% to about 40% of the total nucleic acids in the sample. In some embodiments, minority nucleic acids may be present in amounts ranging from at least about 2% to about 30% of the total nucleic acids in the sample. In some embodiments, minority nucleic acids may be present in amounts ranging from at least about 3% to about 25% of the total nucleic acids in the sample. For example, minority nucleic acids may be present in amounts ranging from about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, or 30% of the total nucleic acids in the sample. In some cases, minority extracellular nucleic acids may make up about 1% to 40% of the total nucleic acid (for example, about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleic acid being minority nucleic acids). In some embodiments, the minority nucleic acids are extracellular DNA.In some embodiments, the minority nucleic acid is extracellular DNA derived from apoptotic tissue. In some embodiments, the minority nucleic acid is extracellular DNA derived from tissue affected by cell proliferation disorder. In some embodiments, the minority nucleic acid is extracellular DNA derived from tumor cells. In some embodiments, the minority nucleic acid is extracellular fetal DNA.
[0070] In another embodiment, the nucleic acids considered to be "major" may constitute, for example, more than 50% to about 99.9% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids may constitute at least about 60% to about 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids may constitute at least about 70% to about 98% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids may constitute at least about 75% to about 97% of the total nucleic acids in the sample. For example, a majority nucleic acid may represent at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acid is extracellular DNA. In some embodiments, the majority nucleic acid is extracellular maternal DNA. In some embodiments, the majority nucleic acid is DNA derived from healthy tissue. In some embodiments, the majority nucleic acid is DNA derived from non-tumor cells.
[0071] In some embodiments, a minority of extracellular nucleic acids are about 500 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 500 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 300 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 300 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 250 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 250 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 200 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 200 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 150 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 150 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 100 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 100 base pairs or less in length). In some embodiments, a minority of extracellular nucleic acids are about 50 base pairs or less in length (for example, about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority nucleic acids are about 50 base pairs or less in length).
[0072] Nucleic acids may be provided for carrying out the methods described herein with or without processing of the sample containing the nucleic acid. In some embodiments, nucleic acids are provided for carrying out the methods described herein after processing of the sample containing the nucleic acid. For example, nucleic acids may be extracted from a sample, isolated, purified, partially purified, or amplified. The term “isolated,” as used herein, means nucleic acids taken out of their original environment (e.g., the natural environment if it is naturally occurring, or the host cell if it is exogenously expressed), and thus modified from their original environment by human intervention (e.g., “by human hands”). The term “isolated nucleic acid,” as used herein, may mean nucleic acids taken out of a subject (e.g., a human subject). Isolated nucleic acids may be provided with less non-nucleic acid components (e.g., proteins, lipids) than are present in the source sample. Compositions containing isolated nucleic acids may contain less than 50% to 99% non-nucleic acid components. A composition containing isolated nucleic acids may contain approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% of non-nucleic acid components. The term “purified,” as used herein, may mean a provided nucleic acid containing less non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount present before the nucleic acid was subjected to the purification procedure. A composition containing purified nucleic acids may contain approximately 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% of other non-nucleic acid components. The term “purified,” as used herein, may mean a provided nucleic acid containing fewer nucleic acid species than the sample source from which the nucleic acid originates. A composition containing purified nucleic acids may contain approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% of other nucleic acid species. For example, fetal nucleic acids can be purified from a mixture containing maternal nucleic acids and fetal nucleic acids.In certain cases, small fragments of fetal nucleic acids (e.g., 30–500 bp fragments) can be purified, or partially purified, from a mixture containing both fetal and maternal nucleic acid fragments. In certain cases, nucleosomes containing smaller fragments of fetal nucleic acids can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acids. In certain cases, nucleic acids from cancer cells can be purified from a mixture containing nucleic acids from cancer cells and nucleic acids from non-cancer cells. In certain cases, nucleosomes containing small fragments of nucleic acids from cancer cells can be purified from a mixture of larger nucleosome complexes containing larger fragments of non-cancer nucleic acids. In some embodiments, nucleic acids are provided for the methods described herein without prior treatment of the sample containing the nucleic acid. For example, nucleic acids can be analyzed directly from a sample without prior extraction, purification, partial purification, and / or amplification.
[0073] In some embodiments, nucleic acids, such as cellular nucleic acids, are sheared or cleaved before, during, or after the methods described herein. The terms “shearing” or “cleavage” generally refer to procedures or conditions that can cleave a nucleic acid molecule (e.g., a nucleic acid template gene molecule or its amplification product) into two (or more) smaller nucleic acid molecules. Such shearing or cleavage may be sequence-specific, base-specific, or non-specific and can be achieved by any of a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, or physical shearing (e.g., physical fragmentation). Sheared or cleaved nucleic acids may have nominal lengths, average lengths, or average lengths of approximately 5 to 10,000 base pairs, 100 to 1,000 base pairs, 100 to 500 base pairs, or approximately 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs.
[0074] Sheared or cleaved nucleic acids can be prepared by preferred methods, including, but not limited to, physical methods (e.g., shearing, e.g., sonication, French press, heating, UV irradiation, etc.), enzymatic processes (e.g., enzymatic cleavage agents (e.g., preferred nucleases, preferred restriction enzymes, preferred methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heating, etc., or combinations thereof), processes described in U.S. Patent Application Publication No. 2005 / 0112590, or combinations thereof. The average length, mean length, or nominal length of the resulting nucleic acid fragments can be controlled by selecting a suitable method for preparing the fragments.
[0075] When used herein, the term “amplified” refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially generates an amplicon nucleic acid or a portion thereof having the same or substantially the same nucleotide sequence as the target nucleic acid. In certain embodiments, the term “amplified” refers to a method comprising a polymerase chain reaction (PCR). In certain cases, the amplification product may contain one or more nucleotides than the nucleotide region being amplified of the nucleic acid template sequence (for example, a primer may contain “extra” nucleotides, such as a transcription start sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplification product containing “extra” nucleotides or nucleotides not corresponding to the nucleotide region being amplified of the nucleic acid template gene molecule).
[0076] Nucleic acids may also be subjected to processes that modify certain nucleotides within the nucleic acid before providing the nucleic acid for the methods described herein. For example, processes that selectively modify the nucleic acid based on the methylation status of nucleotides within the nucleic acid may be applied to the nucleic acid. Furthermore, conditions such as high temperature, ultraviolet light, and X-rays may induce changes in the sequence of nucleic acid molecules. Nucleic acids may be provided in any preferred form useful for sequence analysis.
[0077] Nucleic acid concentration In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched with respect to or relatively to a subpopulation or species of nucleic acids. Subpopulations of nucleic acids may include, for example, fetal nucleic acids, maternal nucleic acids, cancer nucleic acids, patient nucleic acids, nucleic acids containing fragments of a specific length or range of lengths, or nucleic acids derived from a specific genomic region (e.g., a single chromosome, a set of chromosomes and / or a particular chromosomal region). Such enriched samples may be used in conjunction with methods provided herein. Thus, in certain embodiments, the methods of the present technique include a further step of enriching a subpopulation of nucleic acids in a sample, for example, cancer nucleic acids or fetal nucleic acids. In certain embodiments, a method for measuring the proportion of cancer cell nucleic acids or fetal proportions may also be used to enrich cancer nucleic acids or fetal nucleic acids. In certain embodiments, nucleic acids derived from normal tissue (e.g., non-cancerous cells) are selectively removed from the sample (partially, substantially, almost completely, or completely). In certain embodiments, maternal nucleic acids are selectively removed from the sample (partially, substantially, almost completely, or completely). In certain embodiments, quantitative sensitivity can be improved by concentrating a sample for specific low-copy nucleic acids (e.g., cancer nucleic acids or fetal nucleic acids). Methods for concentrating a sample for specific nucleic acid species are described, for example, in U.S. Patent No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of these, including all texts, tables, formulas, and drawings, are incorporated herein by reference.
[0078] In some embodiments, nucleic acids are enriched for a particular target fragment species and / or reference fragment species. In certain embodiments, nucleic acids are enriched for a particular nucleic acid fragment length or range of fragment lengths using one or more length-based separation methods described below. In certain embodiments, nucleic acids are enriched for fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art.
[0079] Non-limiting examples of methods for enriching nucleic acid subpopulations in a sample include: methods that utilize epigenetic differences between nucleic acid species (e.g., a methylation-based method for enriching fetal nucleic acids described in U.S. Patent Application Publication No. 2010 / 0105049, incorporated herein by reference); polymorphic sequencing approaches enhanced by restriction endonucleases (e.g., the method described in U.S. Patent Application Publication No. 2009 / 0317818, incorporated herein by reference); selective enzymatic degradation approaches; large-scale parallel processing signature sequencing (MPSS) approaches; amplification-based approaches (e.g., PCR) (e.g., locus-specific amplification methods, multiplex SNP allele PCR approaches; universal amplification methods); pull-down approaches (e.g., biotinylated ultramer pull-down methods); extension and ligation-based methods (e.g., extension and ligation of molecular inversion probes (MIPs)); and combinations thereof.
[0080] In some embodiments, nucleic acids are enriched for fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based isolation methods described herein. Sequence-based isolation is generally based on nucleotide sequences present in the fragment of interest (e.g., target fragment and / or reference fragment) and substantially absent or present in very small amounts (e.g., less than 5%) of other fragments in the sample. In some embodiments, sequence-based isolation may produce isolated target fragments and / or isolated reference fragments. The isolated target fragments and / or isolated reference fragments are often isolated from the remaining fragments in the nucleic acid sample. In certain embodiments, the isolated target fragments and isolated reference fragments are also isolated from each other (e.g., isolated into separate assay compartments). In certain embodiments, the isolated target fragments and isolated reference fragments are isolated together (e.g., isolated into the same assay compartment). In some embodiments, unbound fragments may be differentially removed, degraded, or digested.
[0081] In some embodiments, a selective nucleic acid capture process is used to separate a target fragment and / or a reference fragment from a nucleic acid sample. Commercially available nucleic acid capture systems include, for example, the Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); the Illumina BEADARRAY platform (Illumina, San Diego, CA); and the Affymetrix GENECHIP platform (Affymetrix, Santa Examples include the Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA) and related platforms. Such methods typically involve hybridization of capture oligonucleotides with some or all of the nucleotide sequence of a target or reference fragment, and may involve the use of solid-phase (e.g., solid-phase arrays) and / or solution-based platforms. Capture oligonucleotides (sometimes referred to as “baits”) may be selected or designed to preferentially hybridize to nucleic acid fragments from selected genomic regions or loci (e.g., one of chromosomes 21, 18, 13, X or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., methods using oligonucleotide arrays) may be used to enrich nucleic acid sequences, genes or regions of interest, from a particular chromosome (e.g., a potentially aneuploid chromosome, a reference chromosome, or another chromosome of interest). Thus, in some embodiments, a nucleic acid sample is enriched as needed by capturing a subset of fragments using, for example, capture oligonucleotides complementary to selected genes in the sample nucleic acid. In certain cases, captured fragments are amplified. For example, a captured fragment containing an adapter may be amplified using a primer complementary to the adapter oligonucleotide to form a collection of amplified fragments indexed according to the adapter sequence. In some embodiments, nucleic acids are enriched for fragments from a selected genomic region (e.g., chromosome, gene) by amplifying one or more regions of interest using oligonucleotides (e.g., PCR primers) complementary to the sequence in the fragment containing the region of interest or a portion thereof.
[0082] In some embodiments, nucleic acids are concentrated for a specific nucleic acid fragment length, length range, or length below or above a specific threshold or cutoff using one or more length-based separation methods. The length of a nucleic acid fragment usually refers to the number of nucleotides in that fragment. The length of a nucleic acid fragment is sometimes referred to as the size of the nucleic acid fragment. In some embodiments, the length-based separation method is performed without measuring the length of individual fragments. In some embodiments, the length-based separation method is performed in conjunction with a method for measuring the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure in which all or part of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on arrays, separation by molecular sieves, separation by gel electrophoresis, separation by column chromatography (e.g., size exclusion columns), and microfluidics-based approaches). In certain cases, length-based separation approaches may include, for example, selective sequence tagging approaches, fragment cyclization, chemical treatments (e.g., formaldehyde, polyethylene glycol (PEG) precipitation), mass spectrometry, and / or size-specific nucleic acid amplification.
[0083] Nucleic acid quantification The amount of nucleic acids in a sample (e.g., concentration, relative amount, absolute amount, copy number, etc.) can be measured. In some embodiments, the amount of minority nucleic acids in the nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain embodiments, the amount of minority nucleic acid species in a sample is referred to as the "minority species ratio." In some embodiments, the "minority species ratio" refers to the ratio of minority nucleic acid species in circulating cell-free nucleic acids in a sample obtained from a subject (e.g., blood sample, serum sample, plasma sample, urine sample).
[0084] The amount of minority nucleic acids in extracellular nucleic acids can be quantified and used in conjunction with the methods provided herein. Therefore, in certain embodiments, the methods described herein include a further step of measuring the amount of minority nucleic acids. The amount of minority nucleic acids in a subject-derived sample can be measured before or after processing for preparing the sample nucleic acid. In certain embodiments, the amount of minority nucleic acids in the sample after processing and preparation of the sample nucleic acid is measured and used for further evaluation. In some embodiments, the outcome includes considering the proportion of minority species in the sample nucleic acid (e.g., adjusting the count, removing the sample, generating or not generating a call).
[0085] The measurement of minority species proportions may be performed before, during, or at any point in time in the methods described herein, or after a particular method described herein (e.g., detection of gene mutations or genetic alterations). For example, in order to perform a gene mutation / gene alteration measurement method with a particular sensitivity or specificity, a minority nucleic acid quantification method may be performed before, during, or after the measurement of gene mutations / gene alterations to identify samples containing minority nucleic acids in amounts greater than or equal to approximately 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, or 25%. In some embodiments, samples measured to contain a certain threshold amount of minority nucleic acids (e.g., about 15% or more; about 4% or more) are further analyzed for, for example, gene mutations / genetic alterations, or the presence or absence of gene mutations / genetic alterations. In certain embodiments, for example, the measurement of gene mutations or genetic alterations is selected only for samples containing a certain threshold amount of minority nucleic acids (e.g., about 15% or more; about 4% or more). (For example, selected and contacted with the patient.)
[0086] In some embodiments, the amount of cancer cell nucleic acid in nucleic acids (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain cases, the amount of cancer cell nucleic acid in a sample is referred to as the "proportion of cancer cell nucleic acid," and may also be referred to as the "cancer ratio" or "tumor ratio." In some embodiments, the "proportion of cancer cell nucleic acid" refers to the proportion of cancer cell nucleic acid in circulating cell-free nucleic acids in a sample obtained from a subject (e.g., blood sample, serum sample, plasma sample, urine sample).
[0087] In some embodiments, the amount of fetal nucleic acid in nucleic acids (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal ratio." In some embodiments, the "fetal ratio" refers to the ratio of fetal nucleic acid to circulating cell-free nucleic acid in a sample obtained from a pregnant woman (e.g., blood sample, serum sample, plasma sample, urine sample). Certain methods described herein or known in the art for measuring the fetal ratio can be used to measure the ratio and / or minority ratio of cancer cell nucleic acids.
[0088] In certain cases, fetal proportions may be measured according to markers specific to male fetuses (e.g., Y-chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), according to the ratio of polymorphic sequence alleles, or according to one or more markers specific to fetal nucleic acids but not specific to maternal nucleic acids (e.g., differential epigenetic biomarkers between mother and fetus (e.g., methylation) or fetal RNA markers in maternal plasma (e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry 53(3):293-296)). Measurement of fetal proportions may sometimes be performed using fetal quantity assays (FQAs), such as those described in U.S. Patent Application Publication No. 2010 / 0105049 (incorporated herein by reference). This type of assay allows for the detection and quantification of fetal nucleic acids in a maternal sample based on the methylation status of the nucleic acids in that sample.
[0089] In certain embodiments, the minority ratio may be measured based on the ratio of alleles of a polymorphic sequence (e.g., a single nucleotide polymorphism (SNP)) using, for example, the method described in U.S. Patent Application Publication No. 2011 / 0224087 (materially incorporated herein by reference). In such a method for measuring fetal ratios, for example, nucleotide sequence reads are obtained for a maternal sample and the fetal ratio is measured by comparing the total number of nucleotide sequence reads that map to a first allele at an informative polymorphic site (e.g., an SNP) in a reference genome with the total number of nucleotide sequence reads that map to a second allele.
[0090] In some embodiments, the minority species ratio may be measured using a method that incorporates information obtained from chromosomal abnormalities, for example, as described in International Patent Application Publication No. WO2014 / 055774 (incorporated herein by reference). In some embodiments, the minority species ratio may be measured using a method that incorporates information obtained from sex chromosomes, for example, as described in U.S. Patent Application Publication Nos. 2013 / 0288244 and U.S. Patent Application Publication Nos. 2013 / 0338933 (each incorporated herein by reference).
[0091] The minority ratio may, in some embodiments, be measured using methods that incorporate information on fragment length (e.g., analysis of the fragment length ratio (FLR), analysis of the fetal ratio statistic (FRS), as described in International Patent Application Publication No. 2013 / 177086 (incorporated herein by reference)). Cell-free fetal nucleic acid fragments are typically shorter than maternal nucleic acid fragments (e.g., see Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Therefore, the fetal ratio may, in some embodiments, be measured by counting fragments below a certain length threshold and comparing their number to, for example, the number of fragments above a certain length threshold and / or the total amount of nucleic acid in the sample. Methods for counting nucleic acid fragments of a specific length are described in further detail in International Patent Application Publication No. WO2013 / 177086.
[0092] The minority proportion may, in some embodiments, be measured according to a partially specific proportion estimation (e.g., as described in International Patent Application Publication No. WO2014 / 205401, incorporated herein by reference). While not bound by theory, the amount of reads from fetal CCF fragments (e.g., fragments of a particular length or length range) often maps to parts (e.g., within the same sample, e.g., within the same sequencing run) with varying frequency. Also, while not bound by theory, certain parts tend to have similar presentations of reads from fetal CCF fragments (e.g., fragments of a particular length or length range) when compared across multiple samples, and this presentation correlates with a partially specific fetal proportion (e.g., the relative amount, percentage, or ratio of CCF fragments originating from the fetus). Partially specific fetal proportion estimates are typically measured according to the relationship between partially specific parameters and their fetal proportions.
[0093] In some embodiments, measuring minority ratios (e.g., ratios of cancer cell nucleic acids; fetal ratios) is not required or necessary to determine the presence or absence of gene mutations or genetic alterations. In some embodiments, determining the presence or absence of gene mutations or genetic alterations does not require distinguishing between sequences of many nucleic acids and sequences of few nucleic acids. In certain embodiments, this is because the sum of contributions from both few and many sequences in a particular chromosome, chromosomal portion, or part thereof is analyzed. In some embodiments, determining the presence or absence of gene mutations or genetic alterations does not rely on speculative sequence information that can distinguish few nucleic acids from many nucleic acids.
[0094] Nucleic acid library In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) prepared, assembled, and / or modified for a particular process, non-limiting examples of which include immobilization onto a solid phase (e.g., a solid support, flow cell, or beads), concentration, amplification, cloning, detection, and / or nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during the sequencing process. The nucleic acid library (e.g., a sequencing library) may be prepared by preferred methods known in the art. The nucleic acid library may be prepared by targeted or untargeted preparation processes.
[0095] In some embodiments, the nucleic acid library is modified to include chemical moieties (e.g., functional groups) configured to immobilize the nucleic acids on a solid support. In some embodiments, the nucleic acid library is modified to include biomolecules (e.g., functional groups) and / or members of binding pairs configured to immobilize the library on a solid support, and non-limiting examples thereof include thyroxine-binding globulins, steroid-binding proteins, antibodies, antigens, haptens, enzymes, lectins, nucleic acids, repressors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and combinations thereof. Some examples of specific binding pairs include, but are not limited to, avidin and biotin moieties; antigenic epitopes and antibodies or their immunologically reactive fragments; antibodies and haptens; digoxigenin moieties and anti-digoxigenin antibodies; fluorescein moieties and anti-fluorescein antibodies; operators and repressors; nucleases and nucleotides; lectins and polysaccharides; steroids and steroid-binding proteins; active compounds and receptors for active compounds; hormones and hormone receptors; enzymes and substrates; immunoglobulins and protein A; oligonucleotides or polynucleotides and their corresponding complementary chains; and combinations thereof.
[0096] In some embodiments, a nucleic acid library is modified to contain one or more polynucleotides of known compositions, non-limiting examples of which include identifiers (e.g., tags, index tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer-binding sites, annealing sites), preferred integration sites (e.g., transposons, viral integration sites), modified nucleotides, or combinations thereof. Polynucleotides of known sequences may be added to preferred positions, for example, the 5' end, the 3' end, or within the nucleic acid sequence. Polynucleotides of known sequences may be the same or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a known 5' sequence may hybridize to a first plurality of oligonucleotides, while a known 3' sequence may hybridize to a second plurality of oligonucleotides. In some embodiments, the nucleic acid library may include chromosome-specific tags, capture sequences, labels, and / or adapters. In some embodiments, the nucleic acid library includes one or more detectable labels. In some embodiments, one or more detectable labels may be incorporated into the nucleic acid library at the 5' end, 3' end, and / or any nucleotide position within the nucleic acid in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes before immobilization on a solid phase.
[0097] In some embodiments, a polynucleotide of a known sequence includes a universal sequence. A universal sequence is a specific nucleotide acid sequence integrated into two or more nucleic acid molecules or two or more subsets of nucleic acid molecules, where the universal sequence is the same for all molecules or subsets of molecules into which it is integrated. Universal sequences are often designed to hybridize to multiple different sequences and / or to amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. Universal primers often include a universal sequence. In some embodiments, an adapter (e.g., a universal adapter) includes a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple nucleic acid species or nucleic acid subsets.
[0098] In certain embodiments of preparing nucleic acid libraries (e.g., in certain sequencing procedures by synthesis), nucleic acids are size-selected and / or fragmented into lengths of several hundred base pairs or less (e.g., in preparation for library construction). In some embodiments, library preparation is performed without fragmentation (e.g., when using cell-free DNA).
[0099] In certain embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego CA). Ligation-based library preparation methods often utilize adapter (e.g., methylated adapter) designs that can incorporate an index sequence (e.g., a sample index sequence that identifies the origin of the sample relative to the nucleic acid sequence) in the initial ligation step, and can often be used to prepare samples for single-read sequencing, paired-end sequencing, and multiplexed sequencing. For example, nucleic acids (e.g., fragmented nucleic acids or cell-free DNA) can have their ends repaired by a fill-in reaction, an exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-end repaired nucleic acid can then be extended by only a single nucleotide complementary to the single nucleotide overhang at the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide.
[0100] In some embodiments, the preparation of a nucleic acid library involves ligating an adapter oligonucleotide (e.g., to a sample nucleic acid, a sample nucleic acid fragment, or a template nucleic acid). The adapter oligonucleotide is often complementary to a flow cell anchor and is sometimes used to immobilize the nucleic acid library to a solid support (e.g., the inner surface of a flow cell). In some embodiments, the adapter oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-ended sequencing primers, paired-ended sequencing primers, multiplexed sequencing primers, etc.) or combinations thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). In some embodiments, the adapter oligonucleotide includes one or more of the following: primer annealing polynucleotides (e.g., for annealing to oligonucleotides attached to a flow cell and / or free amplification primers), index polynucleotides (e.g., sample index sequences for tracking nucleic acids from various samples; also referred to as sample IDs), and barcode polynucleotides (e.g., single-molecule barcodes (SMBs) for tracking individual sample nucleic acid molecules amplified before sequencing; also referred to as molecular barcodes). In some embodiments, the primer annealing component of the adapter oligonucleotide includes one or more universal sequences (e.g., sequences complementary to one or more universal amplification primers). In some embodiments, the index polynucleotide (e.g., sample index; sample ID) is a component of the adapter oligonucleotide. In some embodiments, the index polynucleotide (e.g., sample index; sample ID) is a component of the universal amplification primer sequence.
[0101] In some embodiments, the adapter oligonucleotide is designed to generate a library construct containing one or more of the following when used in combination with an amplification primer (e.g., a universal amplification primer): a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence. In some embodiments, the adapter oligonucleotide is designed to generate a library construct containing an ordered combination of one or more of the following when used in combination with a universal amplification primer: a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, a template sequence (e.g., a sample nucleic acid sequence), a spacer sequence, a second molecular barcode, a third universal sequence, a sample ID, and a fourth universal sequence. In some embodiments, the adapter oligonucleotide is designed to generate a library construct for each strand of a template molecule (e.g., a sample nucleic acid molecule) when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, the adapter oligonucleotide is a double-stranded adapter oligonucleotide.
[0102] An identifier can be a suitable detectable label that is embedded in or attached to a nucleic acid (e.g., a polynucleotide) that enables the detection and / or identification of the nucleic acid containing the identifier. In some embodiments, the identifier is embedded in or attached to the nucleic acid during a sequencing method (e.g., by polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indices or barcodes, radioactive labels (e.g., isotopes), metallic labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore enchanters, dyes, proteins (e.g., enzymes, antibodies or parts thereof, linkers, members of binding pairs), or combinations thereof. In some embodiments, the identifier (e.g., nucleic acid index or barcode) is a unique sequence of nucleotides or nucleotide analogs, a known sequence, and / or an identifiable sequence. In some embodiments, the identifier is six or more consecutive nucleotides. A large number of fluorophores with various different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are used in the methods described herein (e.g., nucleic acid detection methods and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are concatenated to each nucleic acid in the library.Identifier detection and / or quantification may be performed by suitable methods, apparatus or instruments, non-limiting examples thereof, including flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, luminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blotting, mass spectrometry, chromatography, cell fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing apparatus, and combinations thereof.
[0103] In some embodiments, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically involve the simultaneous fragmentation and tagging of DNA in a single-tube reaction (often allowing for the incorporation of platform-specific tags and optional barcodes), using in vitro transposition to prepare a sequencer-ready library.
[0104] In some embodiments, a nucleic acid library or a portion thereof is amplified (e.g., by a PCR-based method). In some embodiments, the sequencing method includes amplification of the nucleic acid library. The nucleic acid library may be amplified before or after immobilization on a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification involves a process of amplifying or increasing the number of existing nucleic acid templates and / or complementary strands (e.g., present in the nucleic acid library) by generating one or more copies of the template and / or complementary strands. Amplification may be carried out by a preferred method. The nucleic acid library may be amplified by thermocycling or isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification is carried out on a solid support (e.g., in a flow cell) on which the nucleic acid library or a portion thereof is immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to an anchor under preferred conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification product is synthesized by extension starting from an immobilized primer. The solid-phase amplification reaction is similar to standard solution-phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support. In some embodiments, modified nucleic acids (e.g., nucleic acids modified by adapter addition) are amplified.
[0105] In some embodiments, solid-phase amplification includes a nucleic acid amplification reaction involving only one type of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes multiple different types of immobilized oligonucleotide primers. In some embodiments, solid-phase amplification may include a nucleic acid amplification reaction involving one type of oligonucleotide primer immobilized on a solid surface and a second different type of oligonucleotide primer in solution. Multiple different types of immobilized or solution-based primers may be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interfacial amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application Publication No. 2013 / 0012399), or combinations thereof.
[0106] Nucleic acid capture In some embodiments, a sample nucleic acid (or sample nucleic acid library) is subjected to a target capture process. Generally, the target capture process is carried out by contacting the sample nucleic acid (or sample nucleic acid library) with a set of probe oligonucleotides under hybridization conditions. The set of probe oligonucleotides (e.g., capture oligonucleotides) generally includes a plurality of probe oligonucleotides having sequences that are complementary or substantially complementary to the sequence in the sample nucleic acid. The plurality of probe oligonucleotides may include about 10 probe oligonucleotide species, about 50 probe oligonucleotide species, about 100 probe oligonucleotide species, about 500 probe oligonucleotide species, about 1,000 probe oligonucleotide species, 2,000 probe oligonucleotide species, 3,000 probe oligonucleotide species, 4,000 probe oligonucleotide species, 5,000 probe oligonucleotide species, 10,000 probe oligonucleotide species, or more. Typically, the first probe oligonucleotide species has a different nucleotide sequence from the second probe oligonucleotide species, and each different species of probe oligonucleotide in a given set has a different nucleotide sequence.
[0107] A probe oligonucleotide typically contains a nucleotide sequence that can hybridize to or anneal to a nucleic acid fragment of interest (e.g., a target fragment) or a portion thereof. Probe oligonucleotides may be naturally occurring or synthetic and may be based on DNA or RNA. A probe oligonucleotide may, for example, be capable of specifically separating a target fragment from other fragments in a nucleic acid sample. The terms “specific” or “specificity,” as used herein, refer to the binding or hybridization of one molecule to another (e.g., an oligonucleotide to a target polynucleotide). “Specific” or “specificity” refers to the recognition, contact, and formation of a stable complex between two molecules, compared to a substantially low recognition, contact, or complex formation between either of the two molecules by another molecule. The terms “anneal” and “hybridize,” as used herein, refer to the formation of a stable complex between two molecules. The terms “probe,” “probe oligonucleotide,” “capture probe,” “capture oligonucleotide,” “capture oligo,” “capture oligo,” “oligo,” or “oligonucleotide” may be used interchangeably throughout this document when referring to probe oligonucleotides.
[0108] Probe oligonucleotides can be designed and synthesized using preferred processes and may be of any length suitable for hybridizing to a target nucleotide sequence and for performing the separation and / or analysis processes described herein. Oligonucleotides can be designed based on a target nucleotide sequence (e.g., a target fragment sequence, a genome sequence, a gene sequence). In some embodiments, oligonucleotides (e.g., probe oligonucleotides) may be about 10 to about 300 nucleotides, about 50 to about 200 nucleotides, about 75 to about 150 nucleotides, about 110 to about 130 nucleotides, or about 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129 nucleotides in length. Oligonucleotides may consist of naturally occurring and / or non-naturally occurring nucleotides (e.g., labeled nucleotides) or mixtures thereof. Oligonucleotides suitable for use in the embodiments described herein can be synthesized and labeled using known methods. Oligonucleotides can be synthesized using an automated synthesizer according to the solid-phase phosphoramidite triester method first reported by Beaucage and Caruthers (1981) Tetrahedron Letts. 22:1859-1862, and / or chemically as described by Needham-VanDevanter et al. (1984) Nucleic Acids Res. 12:6159-6168. Purification of oligonucleotides can be carried out by unmodified acrylamide gel electrophoresis or by anion exchange high-performance liquid chromatography (HPLC), for example, as described by Pearson and Regnier (1983) J. Chrom. 255:137-149.
[0109] In some embodiments, all or part of a probe oligonucleotide sequence (naturally occurring or synthetic) may be substantially complementary to the target sequence or part thereof. “Substantially complementary” means nucleotide sequences that hybridize with each other when referred to herein in relation to sequences. The stringency of the hybridization conditions may be modified to tolerate varying amounts of sequence mismatch. Mutually 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or Beyond that, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more of complementary target sequences and oligonucleotide sequences are included.
[0110] A probe oligonucleotide substantially complementary to a target nucleotide sequence (e.g., a target sequence) or a portion thereof is substantially similar to the complementary strand of the target sequence or its associated portion (e.g., substantially similar to the antisense strand of that nucleic acid). One test to determine whether two nucleotide sequences are substantially similar is to measure the percentage of identical nucleotide sequences shared. "Substantially similar" as referred herein with respect to sequences means that they are substantially similar to each other by 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, This refers to nucleotide sequences that are identical in percentages of 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more.
[0111] Hybridization conditions (e.g., annealing conditions) may be determined and / or adjusted depending on the characteristics of the oligonucleotide used in the assay. The sequence and / or length of the oligonucleotide may sometimes affect hybridization to the target nucleic acid sequence. Depending on the degree of mismatch between the oligonucleotide and the target nucleic acid, low, medium, or high stringency conditions may be used to achieve annealing. As used herein, the term “stringency conditions” refers to the conditions for hybridization and washing. Methods for optimizing the temperature conditions of the hybridization reaction are known in the art and can be found in Current Protocols in Molecular Biology, John Wiley & Sons, NY, 6.3.1–6.3.6 (1989). Aqueous and non-aqueous methods are described in that reference and can be used. A non-limiting example of stringent hybridization conditions is one or more washes in 0.2x SSC, 0.1% SDS at 50°C, following hybridization in 6x sodium chloride / sodium citrate (SSC) at approximately 45°C. Another example of stringent hybridization conditions is one or more washes in 0.2x SSC, 0.1% SDS at 55°C, following hybridization in 6x sodium chloride / sodium citrate (SSC) at approximately 45°C. A further example of stringent hybridization conditions is one or more washes in 0.2x SSC, 0.1% SDS at 60°C, following hybridization in 6x sodium chloride / sodium citrate (SSC) at approximately 45°C. Stringent hybridization conditions often involve hybridization in 6× sodium chloride / sodium citrate (SSC) at approximately 45°C, followed by one or more washes in 0.2× SSC and 0.1% SDS at 65°C. Stringency conditions more often involve 0.5M sodium phosphate and 7% SDS at 65°C, followed by one or more washes in 0.2× SSC and 1% SDS at 65°C.The stringent hybridization temperature can also be altered (i.e., lowered) by adding certain organic solvents, such as formamide. Organic solvents like formamide reduce the thermal stability of double-stranded polynucleotides, allowing hybridization to be performed at lower temperatures while maintaining stringent conditions and extending the lifetime of potentially thermally unstable useful nucleic acids.
[0112] In some embodiments, one or more probe oligonucleotides associate with affinity ligands (e.g., avidin, streptavidin, a capture agent such as an antibody or receptor, a member of a binding pair (e.g., biotin), or an antigen). For example, probe oligonucleotides can be biotinylated so that they can be captured by streptavidin-coated beads.
[0113] In some embodiments, one or more probe oligonucleotides and / or scavengers are effectively bound to a solid support or substrate. The solid support or substrate can be any physically separable solid to which the probe oligonucleotides can directly or indirectly adhere, including, but not limited to, surfaces provided by microarrays and wells, as well as particles, microparticles and nanoparticles such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads). The solid support can be, for example, a chip, column, optical fiber, wipe (wiping paper), filter (e.g., flat surface filter), one or more capillaries, glass and processed or functionalized glass (e.g., porous glass (controlled-pore)). Glass (CPG), quartz, mica, diazotized membranes (paper or nylon), polyformaldehyde, cellulose, cellulose acetate, paper, ceramics, metals, metalloids, semiconductor materials, quantum dots, coated beads or particles, other chromatography materials, magnetic particles; plastics (including acrylic resins, polystyrene, copolymers of styrene or other materials, polybutylene, polyurethane, TEFLON®, polyethylene, polypropylene, polyamide, polyester, polyvinylidene difluoride (PVDF), etc.), polysaccharides, nylon or nitrocellulose, resins, silica Alternatively, silica-based materials (including silicon, silica gel, and modified silicon), Sephadex®, Sepharose®, carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glass, conductive polymers (including polymers such as polypyrrole and polyindole); microstructures or nanostructures (e.g., nucleic acid tiling arrays, nanotubes, nanowires, or surfaces decorated with nanoparticles); or porous surfaces or gels (e.g., methacrylate, acrylamide, sugar polymers, cellulose, silicates, or other fibrous or chain polymers).In some embodiments, the solid support or substrate may be coated with a passively or chemically derivatized coating of any number of materials, including polymers such as dextran, acrylamide, gelatin, or agarose. The beads and / or particles may be free or connected to one another (e.g., sintered). In some embodiments, the solid phase may be an aggregate of particles. In some embodiments, the particles may contain silica, which may contain silicon dioxide. In some embodiments, the silica may be porous, and in certain embodiments, the silica may be non-porous. In some embodiments, the particles further contain an active substance that imparts paramagnetism to the particles. In certain embodiments, the active substance may contain a metal, and in certain embodiments, the active substance may be a metal oxide (e.g., iron or iron oxide, where the iron oxide includes a mixture of Fe2+ and Fe3+). The probe oligonucleotide may be linked to the solid support by covalent or non-covalent interactions, and may be linked to the solid support directly or indirectly (e.g., via an intermediary such as a spacer molecule or biotin). Probe oligonucleotides can be ligated to a solid support before, during, or after nucleic acid capture.
[0114] Nucleic acids modified by means of the addition of adapter sequences described herein may be captured. In some embodiments, unmodified nucleic acids are captured. Nucleic acids may be amplified before and / or after capture by amplification processes such as PCR in some embodiments. The term “captured nucleic acid” typically includes captured nucleic acids and captured and amplified nucleic acids. In some embodiments, captured nucleic acids may be subjected to further capture and amplification. Captured nucleic acids may be sequenced by sequencing processes such as those described herein.
[0115] Detection of copy number variations in captured nucleic acids Methods and processes for classifying the presence or absence of copy number variations (e.g., microduplications, microdeletions) are provided herein. In some embodiments, the determination of the presence or absence of copy number variations is made according to a set of sequence reads. In some embodiments, the determination of the presence or absence of copy number variations is made according to quantitative values of sequence reads for segments and / or subregions described herein. In some embodiments, sequence reads are obtained from circulating cell-free sample nucleic acids from test subjects captured by probe oligonucleotides under hybridization conditions. In some embodiments, the presence or absence of copy number variations is made according to a set of consensus sequences generated from the sequence reads. In some embodiments, the presence or absence of copy number variations is made according to quantitative values of probe coverage. In some embodiments, the determination of the presence or absence of copy number variations is made according to quantitative values of probe coverage for segments and / or subregions described herein. The quantitative values of probe coverage may be quantitative values of sequence reads for each probe oligonucleotide. The quantitative values of probe coverage may be quantitative values of consensus sequences for each probe oligonucleotide. In some embodiments, the presence or absence of copy number mutations is determined according to normalized probe coverage quantification values (e.g., normalized probe coverage quantification values for sequence reads for each probe oligonucleotide; normalized probe coverage quantification values for consensus sequences for each probe oligonucleotide). In some embodiments, the determination of the presence or absence of copy number mutations includes a segmentation process. In some embodiments, the determination of the presence or absence of copy number mutations includes a filtering process.
[0116] In some embodiments, the determination of the presence or absence of copy number variation is based on probe coverage quantification or normalized probe coverage quantification. In some embodiments, "based on" may include other factors (e.g., segments, filtered segments, measurement or estimation of copy number, measurement or estimation of increase or decrease of copy number, measurement or estimation of filtered copy number, measurement or estimation of increase or decrease of filtered copy number). In some embodiments, the presence or absence of copy number variation may be determined according to probe coverage quantification or normalized probe coverage quantification for a single probe oligonucleotide. In some embodiments, the presence or absence of copy number variation may be determined according to probe coverage quantification or normalized probe coverage quantification for multiple probe oligonucleotides.
[0117] In some embodiments, the sample nucleic acid is captured by a probe oligonucleotide. Typically, in such embodiments, the sample nucleic acid is brought into contact with the probe oligonucleotide under hybridization conditions. The sample nucleic acid may contain (or consist of) a sample polynucleotide, and the probe oligonucleotide may contain a probe polynucleotide complementary to the sample polynucleotide in the sample nucleic acid. In some embodiments, the probe polynucleotide is complementary to the sequence in the target subchromosome region, segment, and / or subregion as described herein. In some embodiments, the stringency of the hybridization conditions allows only probe polynucleotides with 100% complementarity (i.e., no mismatches) to hybridize to the sample nucleic acid. In some embodiments, the stringency of the hybridization conditions allows probe polynucleotides with one or two mismatches to hybridize to the sample nucleic acid.
[0118] In some embodiments, sequence reads are mapped to a reference genome portion. Certain methods for mapping sequence reads to a reference genome portion are described herein. In some embodiments, the genome portions are of a predetermined length. In some embodiments, the genome portions are of equal length. In some embodiments, the genome portions are about 50 kilobases long. In some embodiments, at least two genome portions are of unequal length. In some embodiments, the genome portions do not overlap. In some embodiments, the 3' end of a genome portion is adjacent to the 5' end of each adjacent downstream genome portion. In some embodiments, at least two genome portions overlap.
[0119] In some embodiments, sequence reads that map to a reference genome match the probe sequence and are identified as on-target reads. In some embodiments, the methods herein include the step of identifying on-target reads. In some embodiments, a read is identified as on-target when it aligns with a genomic region corresponding to a probe oligonucleotide sequence. As described in further detail herein, probe oligonucleotide sequences often contain nucleotide sequences that align (i.e., correspond) to a specific region of the genome (e.g., a reference genome) and correspond to a specific genomic sequence of interest (e.g., a sequence in a target subchromosome region, segment, and / or subregion as described herein). A read that aligns with the genomic region to which the probe oligonucleotide aligns is considered an on-target read. In some embodiments, a sequence read may be considered on-target when the entire length of the read aligns with the genomic region to which the probe oligonucleotide aligns. In some embodiments, a read is identified as on-target when a portion of the read aligns with a genomic region corresponding to a probe oligonucleotide sequence, and a portion of the read aligns within a genomic region adjacent to the genomic region corresponding to the probe oligonucleotide sequence. Generally, in such cases, the read aligns with a contiguous genomic sequence that includes 1) a portion of the genomic region corresponding to the probe oligonucleotide sequence, and 2) a genomic region adjacent to the genomic region corresponding to the probe oligonucleotide sequence. The latter genomic region may be located upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence.For example, a sequence read may be considered on-target when a portion of the read having a genomic region corresponding to a probe oligonucleotide sequence (e.g., at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%) and the remainder of the read align with a genomic sequence immediately upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence. In some embodiments, a sequence read may be considered on-target when a portion of the read does not align with the probe sequence, but the entire read length aligns with a genomic sequence immediately upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence.
[0120] A sequence containing a probe sequence (i.e., the genomic sequence corresponding to the probe sequence) and further genomic sequences upstream and / or downstream of the probe sequence may be referred to as a padded probe sequence. A collection of padded probe sequences may be referred to as a padded panel. In some embodiments, a padded probe sequence includes at least one nucleotide of a genomic sequence immediately upstream and / or downstream of the genomic sequence corresponding to the probe sequence. For example, a padded probe sequence may include at least about 5, 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, or 1000 nucleotides of a genomic sequence immediately upstream and / or downstream of the genomic sequence corresponding to that probe sequence. In some embodiments, a padded probe sequence includes a 250-nucleotide genomic sequence immediately upstream and a 250-nucleotide genomic sequence immediately downstream of the genomic sequence corresponding to the probe sequence.
[0121] Probe oligonucleotide sequences may be stored in a database as a sequence panel. In some embodiments, reads are directly aligned to probe oligonucleotide sequences (e.g., probe oligonucleotide sequences stored in a table or database with or without adjacent genomic region sequences, as described above), and such reads are identified as on-target reads. For example, sequence reads may be aligned to a sequence panel in a database without first mapping to a reference genome. In some embodiments, a sequence read may be considered on-target when its entire length is aligned to the probe sequence. In some embodiments, sequence reads are directly aligned to a padded probe sequence, as described above. For example, in some embodiments, a sequence read may be considered on-target if a portion of the read (e.g., at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%) aligns with the probe sequence, and the remainder aligns with a genomic sequence immediately upstream or downstream of the probe sequence. In some embodiments, a sequence read may be considered on-target if a portion of the read does not align with the probe sequence, and the entire read length aligns with a genomic sequence immediately upstream or downstream of the probe sequence.
[0122] In some embodiments, the consensus sequence is generated from sequence reads. In some embodiments, the consensus sequence is generated from sequence reads identified as "on-target" reads. Generally, consensus is generated by breaking down a set of sequence reads (e.g., reads within a group of reads) to generate a single nucleotide sequence corresponding to a unique nucleic acid molecule in the sample from which the sequence reads were generated. The consensus sequence can be generated from a group of reads by any preferred method, such as linear or nonlinear methods for consensus generation derived from digital communications theory, information theory, or bioinformatics (e.g., averaging, voting, statistical, dynamic programming, maximum posterior probability or maximum likelihood detection, Bayesian methods, hidden Markov methods, or support vector machine methods).
[0123] In some embodiments, the presence or absence of copy number variation is determined according to probe coverage quantification values (e.g., probe coverage quantification values for segments and / or subregions described herein; probe coverage quantification values for sequences in segments and / or subregions described herein). Probe coverage generally refers to the quantification value of sequence reads or consensus sequences mapped to each nucleotide position in a probe oligonucleotide. In some embodiments, measuring probe coverage quantification values involves measuring the number of sequence reads that map to each nucleotide position in a probe oligonucleotide. Sequence reads may be shorter than the probe oligonucleotide and / or may partially overlap with the probe oligonucleotide sequence. Therefore, the quantification value of sequence reads mapped to each nucleotide in the probe may vary depending on the length of the probe oligonucleotide. Therefore, in some embodiments, measuring probe coverage quantification values involves measuring quantile estimates of the population of sequence reads mapped to each nucleotide position in the probe. Examples of quantile estimates include median, mean, mode, range, etc. In some embodiments, measuring probe coverage quantification includes measuring the median number of sequence reads mapped to each nucleotide position in the probe. In some embodiments, the median number of sequence reads mapped to each nucleotide position for each probe oligonucleotide is the probe coverage quantification for each probe oligonucleotide. In some embodiments, measuring probe coverage quantification includes measuring the number of consensus sequences mapped to each nucleotide position in the probe oligonucleotide. The consensus sequences may be shorter than the probe oligonucleotide and / or may partially overlap with the probe oligonucleotide sequence. Therefore, the quantification of consensus sequences mapped to each nucleotide in the probe may vary depending on the length of the probe oligonucleotide.Therefore, in some embodiments, measuring probe coverage quantification involves measuring the median number of consensus sequences mapped to each nucleotide position in the probe.
[0124] In some embodiments, the presence or absence of copy number variation is determined according to normalized probe coverage quantification values. Probe coverage quantification values may be normalized using a preferred normalization process, such as the normalization processes described herein. In some embodiments, normalization includes scaling the probe coverage quantification values for each probe oligonucleotide to the test sample. Scaling the probe coverage quantification values for each probe oligonucleotide generates a scaled probe coverage quantification value for each probe oligonucleotide. In some embodiments, the probe coverage quantification value for each probe is scaled according to the median of the probe coverage quantification values for all probe oligonucleotides to the test sample. For example, the probe coverage quantification value for each probe oligonucleotide may be divided by the median of the probe coverage quantification values.
[0125] In some embodiments, normalization includes normalizing the probe coverage quantification value according to the guanine-cytosine (CG) content for each probe oligonucleotide in the test sample. In some embodiments, normalization includes normalizing the scaled probe coverage quantification value according to the guanine-cytosine (CG) content for each probe oligonucleotide in the test sample. Normalizing the probe coverage quantification value according to the GC content for each probe oligonucleotide generates a GC-normalized probe coverage quantification value for each probe oligonucleotide. In some embodiments, the probe coverage quantification value is normalized by LOESS normalization. LOESS normalization (e.g., GC LOESS) is described in further detail herein.
[0126] In some embodiments, normalization includes normalizing the probe coverage quantification for the test sample according to the probe coverage quantification obtained from the reference sample. The reference sample may include a sample classified as having no copy number variation. In some embodiments, the reference sample consists of a sample classified as having no copy number variation. Therefore, in some embodiments, the reference sample includes or consists of a sample that is euploid for each chromosome and each chromosomal region being tested. The reference sample may be derived from a human subject. In some embodiments, the reference sample is derived from a female subject. In some embodiments, the reference sample is derived from a male subject. In some embodiments, the reference sample is derived from both male and female subjects. The reference sample may include a sample from one subject or a sample from multiple subjects. The reference sample may include one reference sample, and often includes multiple samples. For example, a reference sample may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more samples.
[0127] In some embodiments, the probe coverage quantification for a test sample is normalized according to the probe coverage quantification obtained from a reference sample. In some embodiments, the scaled probe coverage quantification for a test sample is normalized according to the probe coverage quantification obtained from a reference sample. In some embodiments, the GC-normalized probe coverage quantification for a test sample is normalized according to the probe coverage quantification obtained from a reference sample. In some embodiments, the probe coverage quantification for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, the scaled probe coverage quantification for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, the GC-normalized probe coverage quantification for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. The median probe coverage is often measured according to the probe coverage quantification for the same probe across multiple reference samples. In some embodiments, the median probe coverage is measured according to a normalized (e.g., GC-normalized) probe coverage quantification for the same probe across multiple reference samples. Normalizing the probe coverage quantification (or scaled probe coverage quantification or GC-normalized probe coverage quantification) for each probe oligonucleotide according to the probe coverage quantification obtained from the reference samples (e.g., median probe coverage) generates a reference-normalized probe coverage quantification for each probe oligonucleotide relative to the test sample.
[0128] In some embodiments, normalization according to the median probe coverage (e.g., the median probe coverage for each probe oligonucleotide obtained from a reference sample) includes dividing each probe coverage quantification for each probe oligonucleotide (i.e., for the test sample) by the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, normalization according to the median probe coverage (e.g., the median probe coverage for each probe oligonucleotide obtained from a reference sample) includes dividing each scaled probe coverage quantification for each probe oligonucleotide (i.e., for the test sample) by the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, normalization according to the median probe coverage (e.g., the median probe coverage for each probe oligonucleotide obtained from a reference sample) includes dividing each GC-normalized probe coverage quantification for each probe oligonucleotide (i.e., for the test sample) by the median probe coverage for each probe oligonucleotide obtained from a reference sample. In such embodiments, a ratio for each probe oligonucleotide is generated by normalizing according to the median probe coverage.
[0129] In some embodiments, the probe coverage quantitative values are logarithmically transformed. For example, the probe coverage quantitative values normalized to the reference sample for each probe oligonucleotide can be logarithmically transformed. By logarithmically transforming the probe coverage quantitative values normalized to the reference sample for each probe oligonucleotide, a logarithmically transformed, reference-sample-normalized probe coverage quantitative value for each probe oligonucleotide is generated. In certain embodiments, the ratio for each probe oligonucleotide is logarithmically transformed. By logarithmically transforming the ratio for each probe oligonucleotide, a logarithmically transformed ratio for each probe oligonucleotide is generated. In some embodiments, the logarithmic transformation is a log2 transformation. Therefore, in some embodiments, a log2 transformed, reference-sample-normalized probe coverage quantitative value for each probe oligonucleotide is generated. In some embodiments, a log2 ratio for each probe oligonucleotide is generated. In a particular case, the log2 ratio of the probe coverage quantitative values is, for example, given by equation A:
number
[0130] In the formula, “test coverage” refers to the probe coverage quantified value for the probe oligonucleotide relative to the test sample (e.g., scaled probe coverage quantified value, normalized probe coverage quantified value); “normal coverage” refers to the probe coverage quantified value for the probe oligonucleotide obtained from the reference sample (e.g., median probe coverage); and CN is the copy number increase or copy number decrease for the segment represented by the probe oligonucleotide (i.e., the segment containing a sequence identical or substantially identical to the probe oligonucleotide sequence).
[0131] The normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to any normalized probe coverage quantification or any preferred variation thereof as described herein. For example, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to a scaled probe coverage quantification for a probe oligonucleotide relative to a test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to a GC-normalized probe coverage quantification for a probe oligonucleotide relative to a test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to a reference sample-normalized probe coverage quantification for a probe oligonucleotide relative to a test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to a ratio of a probe oligonucleotide relative to a test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to a logarithmically transformed reference sample-normalized probe coverage quantification for a probe oligonucleotide relative to a test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to the logarithmically transformed ratio of the probe oligonucleotide relative to the test sample. In certain cases, the normalized probe coverage quantification for a probe oligonucleotide relative to a test sample may refer to the log2 transformed, reference-normalized probe coverage quantification for the probe oligonucleotide relative to the test sample.In certain cases, the normalized probe coverage quantification value for the probe oligonucleotide relative to the test sample may refer to the log2 ratio of the probe oligonucleotide relative to the test sample.
[0132] In some embodiments, the segmentation process is applied to identify segments (e.g., segments spanning copy number variations). Any suitable segmentation process may be used, including, but not limited to, the circular binary segmentation (CBS) process. Other processes may be used instead of or in addition to CBS, non-limiting examples of which include wavelet segmentation (e.g., Haar wavelet segmentation), Fourier transform, sliding window z-scores, and Markov chain models.
[0133] In some embodiments, the segmentation process is applied to identify segments according to probe coverage quantification values for each probe oligonucleotide. In some embodiments, the segmentation process is applied to identify segments according to normalized probe coverage quantification values for each probe oligonucleotide. A segment may contain multiple probe oligonucleotides (i.e., multiple probe oligonucleotides having probe coverage quantification values that suggest an increase or decrease in copy number variation). The segmentation process may provide start and end locations for each segment (e.g., start and end locations according to genomic coordinates; start and end locations according to probe index), copy number variation quantification values for the segment, and, if necessary, a measure of confidence for the segment. In some embodiments, the location for each end of each segment (e.g., location according to probe index) and probe coverage quantification values are provided for each segment. In some embodiments, the location for each end of each segment (e.g., location according to probe index) and normalized probe coverage quantification values are provided for each segment. In some embodiments, one or more genes overlapping with each segment are identified.
[0134] In some embodiments, the copy number of each segment is determined or estimated according to the probe coverage quantification value associated with each segment. In some embodiments, the copy number of each segment is determined or estimated according to the normalized probe coverage quantification value associated with each segment. The determination or estimation of the copy number of each segment provides a copy number (CN) increase or copy number (CN) decrease for each segment. In some embodiments, the copy number (CN) increase or copy number (CN) decrease for each segment is determined or estimated according to the transformation of the segment median coverage for each segment. Thus, in a particular case, the segment median coverage is determined according to the probe coverage quantification value for the probe oligonucleotide in the segment. In some embodiments, the copy number (CN) increase or copy number (CN) decrease for a segment is determined or estimated according to the transformation of the segment median coverage log2 ratio for each segment. Thus, in a particular case, the median log2 ratio is determined according to the probe coverage quantification value for the probe oligonucleotide in the segment. In other words, the median log2 ratio to the probe oligonucleotide in a segment is used to determine or estimate copy number (CN) increase or decrease in the segment. For example, copy number increase or decrease in a segment is given by equation B: CN=2 * (2 (セグメント.中央値.log2比) -1) Equation B This can be determined or estimated according to the formula (wherein CN is a copy number increase or copy number decrease for each segment).
[0135] In some embodiments, segments are filtered (e.g., removed from consideration). Segments may be filtered according to a probe coverage quantifier associated with the segment, a normalized probe coverage quantifier associated with the segment, and one or more copy number increases or decreases relative to the segment. Typically, segment filtering provides a set of filtered and retained segments. Segments are often paired with a corresponding copy number quantifier, and segments where the absolute value of the corresponding copy number quantifier is between 0 and approximately 1 (for duplicate candidates) or between 0 and approximately 0.9 (for deletion candidates) are often filtered and removed as part of the noise reduction filtering process. In some embodiments, segments with a copy number increase of 1 or more (for duplicate candidates) may be retained as filtered segments. For example, segments with copy number increases of 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more may be retained as filtered segments. In some embodiments, segments with a copy number reduction of 0.9 or greater (in the case of deletion candidates) may be retained as filtered segments. Since the copy number quantity for deletion candidates typically falls below zero, "copy number reduction of 0.9 or greater" corresponds to the absolute value of the copy number quantity for deletion candidates. Therefore, for example, segments with copy number reductions of 0.9, 1.0, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 may be retained as filtered segments. In other words, segments with copy number quantities of -0.9, -1.0, -1.2, -1.3, -1.4, -1.5, -1.6, -1.7, -1.8, -1.9, or -2 may be retained as filtered segments.
[0136] Nucleic acid sequencing and processing The methods provided herein typically involve sequencing and analysis of nucleic acids. In some embodiments, nucleic acids are sequenced, and the sequencing product (e.g., a collection of sequence reads) is processed before or concurrently with the analysis of the sequenced nucleic acid. For example, sequence reads may be processed according to one or more of the following steps: alignment, mapping, part filtering, part selection, counting, normalization, weighting, profile generation, etc., and any combination thereof. Certain processing steps may be performed in any order, and certain processing steps may be repeated. For example, after part filtering, the sequence read count may be normalized, and in certain embodiments, after the sequence read count has been normalized, part filtering may be performed. In some embodiments, after the part filtering step, the normalization of the sequence read count is followed by a step of further part filtering. Certain sequencing methods and processing steps are described in more detail below.
[0137] Sequence determination In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain cases, a complete or substantially complete sequence may be obtained, and sometimes only a partial sequence may be obtained. Nucleic acid sequencing typically produces a collection of sequence reads. As used herein, “reads” (e.g., “a read,” “sequence read”) are short nucleotide sequences produced by any sequencing process described herein or known in the art. Reads may be generated from one end of a nucleic acid fragment (“single-ended read”), or from both ends of a nucleic acid fragment (e.g., paired-ended read, double-ended read).
[0138] The length of a sequence read is often related to the specific sequencing technique. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). For example, nanopore sequencing can provide sequence reads that can vary in size from tens, hundreds to thousands of base pairs. In some embodiments, a sequence read is an average, median, mean length, or absolute length of approximately 15 bp to approximately 900 bp. In certain embodiments, a sequence read is an average, median, mean length, or absolute length of approximately 1000 bp or greater. In some embodiments, a sequence read is an average, median, mean length, or absolute length of approximately 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 bp or greater. In some embodiments, a sequence read is an average, median, mean length, or absolute length of approximately 100 bp to approximately 200 bp. In some embodiments, the sequence reads are mean, median, average length, or absolute length of approximately 140 bp to approximately 160 bp. For example, the sequence reads may be mean, median, average length, or absolute length of approximately 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, or 160 bp.
[0139] In some embodiments, the nominal length, average length, mean length, or absolute length of a single-ended lead may be approximately 10 consecutive nucleotides to approximately 250 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 200 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 150 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 125 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 100 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 75 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 60 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 50 or more consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 40 or more consecutive nucleotides, and may be approximately 15 consecutive nucleotides or approximately 36 or more consecutive nucleotides. In certain embodiments, the nominal length, average length, mean length, or absolute length of a single-ended read is about 20 to about 30 base pairs or about 24 to about 28 base pairs. In certain embodiments, the nominal length, average length, mean length, or absolute length of a single-ended read is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, or about 29 base pairs or more. In certain embodiments, the nominal length, average length, mean length, or absolute length of a single-ended read is about 20 to about 200 base pairs, about 100 to about 200 base pairs or about 140 to about 160 base pairs. In a particular embodiment, the nominal length, average length, mean length, or absolute length of a single-ended read is approximately 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or approximately 200 base pairs or longer.In certain embodiments, the nominal length, average length, mean length, or absolute length of a paired-end read may be approximately 10 to approximately 25 consecutive nucleotides or more (e.g., approximately 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length or more), approximately 15 to approximately 20 consecutive nucleotides or more, or approximately 17 or approximately 18 consecutive nucleotides. In a particular embodiment, the nominal length, average length, mean length, or absolute length of a paired-end read is approximately 25 consecutive nucleotides to approximately 400 consecutive nucleotides or more (e.g., approximately 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 nucleotide lengths or more), approximately 50 consecutive nucleotides to approximately 350 The number of consecutive nucleotides may be several nucleotides or more, approximately 100 to approximately 325 consecutive nucleotides, approximately 150 to approximately 325 consecutive nucleotides, approximately 200 to approximately 325 consecutive nucleotides, approximately 275 to approximately 310 consecutive nucleotides, approximately 100 to approximately 200 consecutive nucleotides, approximately 100 to approximately 175 consecutive nucleotides, approximately 125 to approximately 175 consecutive nucleotides, and approximately 140 to approximately 160 consecutive nucleotides. In certain embodiments, the nominal length, average length, mean length, or absolute length of a paired-end read is approximately 150 consecutive nucleotides, and may be 150 consecutive nucleotides.
[0140] In some embodiments, the nucleotide sequence reads obtained from a sample are partial nucleotide sequence reads. As used herein, “partial nucleotide sequence read” refers to a sequence read of any length that has incomplete sequence information, also known as sequence ambiguity. Partial nucleotide sequence reads may lack information about the identity and / or position or order of nucleic acid bases. Partial nucleotide sequence reads generally do not include sequence reads whose incomplete sequence information (or fewer than all of those bases have been sequenced or are determined) stems from careless or unintentional sequencing errors. Such sequencing errors may be specific to a particular sequencing process and include, for example, incorrect calls to nucleic acid base identity and missing or extra nucleic acid bases. Therefore, certain information about the sequence of partial nucleotide sequence reads is often deliberately excluded herein. That is, sequence information relating to fewer than all nucleic acid bases, or sequence information that can be otherwise characterized as or could be a sequencing error, is deliberately obtained. In some embodiments, a partial nucleotide sequence read may extend to a portion of a nucleic acid fragment. In some embodiments, a partial nucleotide sequence read may extend to the entire length of a nucleic acid fragment. Partial nucleotide sequence reads are described, for example, in International Patent Application Publication No. WO2013 / 052907, the entirety of which, including all text, tables, formulas and drawings, is incorporated herein by reference.
[0141] A read is generally a presentation of a nucleotide sequence in a physical nucleic acid. For example, in a read containing the sequence of an ATGC description, in the physical nucleic acid, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide. Sequence reads obtained from a sample derived from a subject may be reads from a mixture of few nucleic acids and many nucleic acids. For example, a sequence read obtained from the blood of a cancer patient may be a read from a mixture of cancer nucleic acids and non-cancer nucleic acids. In another example, a sequence read obtained from the blood of a pregnant woman may be a read from a mixture of fetal nucleic acids and maternal nucleic acids. A mixture of relatively short reads may be converted by the processes described herein to presentations of genomic nucleic acids present in the subject and / or present in a tumor or fetus. In certain cases, a mixture of relatively short reads may be converted, for example, to presentations of copy number variations, gene mutations / genetic alterations or aneuploidy. In one example, a read of a mixture of cancer nucleic acids and non-cancerous nucleic acids may be converted into a presentation of a composite chromosome or a portion thereof containing chromosomal features of one or both cancer cells and non-cancerous cells. In another example, a read of a mixture of maternal nucleic acids and fetal nucleic acids may be converted into a presentation of a composite chromosome or a portion thereof containing chromosomal features of one or both maternal and fetal cells.
[0142] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from cancer patients include nucleic acid fragments originating from normal cells (i.e., non-cancerous fragments) and nucleic acid fragments originating from cancer cells (i.e., cancerous fragments). Sequence reads derived from CCF fragments originating from normal cells (i.e., non-cancerous cells) are referred to herein as “non-cancerous reads.” Sequence reads derived from CCF fragments originating from cancerous cells are referred to herein as “cancerous reads.” CCF fragments from which non-cancerous reads are obtained may be referred to herein as non-cancerous templates, and CCF fragments from which cancerous reads are obtained may be referred to herein as cancerous templates.
[0143] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from pregnant women include nucleic acid fragments originating from fetal cells (i.e., fetal fragments) and nucleic acid fragments originating from maternal cells (i.e., maternal fragments). Sequence reads derived from fetal-origin CCF fragments are referred to herein as "fetal reads." Sequence reads derived from CCF fragments originating from the genome of a pregnant woman with a fetus (e.g., the mother) are referred to herein as "maternal reads." CCF fragments from which fetal reads are obtained are referred to herein as fetal templates, and CCF fragments from which maternal reads are obtained are referred to herein as maternal templates.
[0144] In certain embodiments, “obtaining” nucleic acid sequence reads of a sample from a subject and / or “obtaining” nucleic acid sequence reads of a biological specimen from one or more reference persons may include directly sequencing the nucleic acid to obtain sequence information. In some embodiments, “obtaining” may include receiving sequence information obtained directly from the nucleic acid by another.
[0145] In some embodiments, some or all nucleic acids in a sample are enriched and / or amplified (e.g., nonspecifically, e.g., by a PCR-based method) before or during sequencing. In certain embodiments, specific nucleic acid species or subsets in a sample are enriched and / or amplified before or during sequencing. In some embodiments, species or subsets of a pre-selected nucleic acid pool are randomly sequenced. In some embodiments, nucleic acids in a sample are not enriched and / or amplified before or during sequencing.
[0146] In some embodiments, a representative portion of the genome is sequenced, which is sometimes referred to as “coverage” or “double coverage.” For example, 1x coverage suggests that approximately 100% of the nucleotide sequences of that genome are represented by reads. In some cases, double coverage is referred to as “sequencing depth” (and is directly proportional to “sequencing depth”). In some embodiments, “double coverage” is a relative term referring to a prior sequencing run. For example, a second sequencing run may have twice as little coverage as a first sequencing run. In some embodiments, the genome is sequenced redundantly, where a given genomic region may be covered by two or more reads or overlapping reads (e.g., “double coverage” greater than 1, e.g., 2x coverage). In some embodiments, the genome (e.g., whole genome) is sequenced with coverage of approximately 0.01x to approximately 100x, approximately 0.1x to approximately 20x, or approximately 0.1x to approximately 1x (e.g., approximately 0.015, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90x, or more). In some embodiments, a specific portion of the genome (e.g., a genomic portion by a targeting method and / or a probe-based method) is sequenced, and the double coverage value usually refers to a portion of that specific genomic portion that has been sequenced (i.e., the double coverage value does not refer to the entire genome). In some cases, a specific genomic portion is sequenced with 1,000x coverage or more. For example, a specific genomic portion may be sequenced with 2,000x, 5,000x, 10,000x, 20,000x, 30,000x, 40,000x, or 50,000x coverage. In some embodiments, sequencing is performed with coverage of about 1,000x to about 100,000x.In some embodiments, sequencing is performed with coverage of approximately 10,000x to approximately 70,000x. In some embodiments, sequencing is performed with coverage of approximately 20,000x to approximately 60,000x. In some embodiments, sequencing is performed with coverage of approximately 30,000x to approximately 50,000x.
[0147] In some embodiments, a single nucleic acid sample from a single individual is sequenced. In a particular embodiment, nucleic acids from each of two or more samples are sequenced, where the samples are from one individual or from different individuals. In a particular embodiment, nucleic acid samples from two or more biological samples are pooled, where each biological sample is from one individual or from two or more individuals, and the pool is sequenced. In the latter embodiment, the nucleic acid samples from each biological sample are often identified by one or more unique identifiers.
[0148] In some embodiments, the sequencing method uses identifiers that enable the multiplexing of sequencing reactions in the sequencing process. The more unique identifiers there are, the more samples and / or chromosomes can be multiplexed in the sequencing process, for example, for detection. The sequencing process can be carried out using any suitable number of unique identifiers (e.g., 4, 8, 12, 24, 48, 96 or more).
[0149] The sequencing process may utilize a solid phase, which may include a flow cell to which nucleic acids from the library may be attached, through which reagents may flow, and which may come into contact with the attached nucleic acids. The flow cell may have flow cell lanes, and the use of identifiers may facilitate the analysis of several samples in each lane. The flow cell is often a solid support that can be configured to hold a reagent solution on a bound analyte and / or to pass the reagent solution over the bound analyte in an orderly manner. The flow cell is often planar in shape, optically transparent, generally on a millimeter or sub-millimeter scale, and often has channels or lanes through which analyte / reagent interactions occur. In some embodiments, the number of samples analyzed in a given flow cell lane depends on the number of unique identifiers used during library preparation and / or probe design. Multiplexing using 12 identifiers, for example, allows for the simultaneous analysis of 96 samples (e.g., equal to the number of wells in a 96-well microwell plate) in an 8-lane flow cell. Similarly, multiplexing using 48 identifiers allows for the simultaneous analysis of 384 samples (e.g., the number of wells in a 384-well microwell plate) in an 8-lane flow cell. Non-limiting examples of commercially available multiplex sequencing kits include Illumina's multiplex sample preparation oligonucleotide kits and multiplex sequencing primers and PhiX control kits (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).
[0150] Any preferred method for sequencing nucleic acids may be used, non-limiting examples of which include Maxim & Gilbert, chain termination, synthesis sequencing, ligation sequencing, mass spectrometry sequencing, microscopy-based methods, or combinations thereof. In some embodiments, first-generation techniques, such as Sanger sequencing methods including automated Sanger sequencing including microfluidics Sanger sequencing, may be used in the methods provided herein. In some embodiments, sequencing techniques including the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)) may be used. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally require a clone-amplified DNA template or a single DNA molecule to be sequenced in a large-scale parallel processing format, sometimes in a flow cell. Next-generation (e.g., second and third-generation) sequencing methods capable of sequencing DNA in a large-scale parallel processing format may be used for the methods described herein and are collectively referred to herein as “Large-Scale Parallel Processing Sequences” (MPS). In some embodiments, MPS sequencing uses a targeted approach, where a specific chromosome, gene, or region of interest is sequenced. In certain embodiments, a non-targeted approach is used, in which most or all nucleic acids in a sample are randomly sequenced, amplified, and / or captured.
[0151] In some embodiments, targeted enrichment, amplification, and / or sequencing approaches are used. Targeted approaches often isolate, select, and / or enrich a subset of nucleic acids in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is used to target (e.g., hybridize) one or more sets of nucleic acids in a sample. Sequence-specific oligonucleotides and / or primers are often selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more chromosomes, genes, exons, introns, and / or regulatory regions of interest. Any preferred method or combination of methods may be used for enrichment, amplification, and / or sequencing of one or more targeted subsets of nucleic acids. In some embodiments, the targeted sequences are isolated and / or enriched by capture to a solid phase (e.g., flow cell, beads) using one or more sequence-specific anchors. In some embodiments, the targeted sequence is enriched and / or amplified by a polymerase-based method (e.g., a PCR-based method, or extension based on any suitable polymerase) using sequence-specific primers and / or primer sets. Sequence-specific anchors can often be used as sequence-specific primers.
[0152] MPS sequencing sometimes utilizes sequencing by synthesis and certain imaging processes. Nucleic acid sequencing techniques that may be used in the methods described herein include synthetic sequencing and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego, CA)). This technique allows for the parallel sequencing of millions of nucleic acid (e.g., DNA) fragments. One example of this type of sequencing technique involves a flow cell with an optically transparent slide having eight separate lanes on a surface to which oligonucleotide anchors (e.g., adapter primers) are bound.
[0153] Synthetic sequencing is typically performed by repeatedly adding nucleotides to a primer or existing nucleic acid chain in a template-specific manner (e.g., covalent addition). Each repeated addition of nucleotides is detected, and the process is repeated multiple times until the sequence of the nucleic acid chain is obtained. The length of the resulting sequence depends in part on the number of addition and detection steps performed. In some embodiments of synthetic sequencing, one, two, three or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a single nucleotide addition. Nucleotides can be added by any preferred method (e.g., enzymatically or chemically). For example, in some embodiments, polymerases or ligases add nucleotides to the primer or existing nucleic acid chain in a template-specific manner. In some embodiments of synthetic sequencing, different types of nucleotides, nucleotide analogs, and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequencing by synthesis includes cleavage (e.g., cleavage and removal of identifiers) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by preferred methods described herein or known in the art, non-limiting examples of which include any preferred imaging device, preferred camera, digital camera, CCD (charge-coupled device) based imaging device (e.g., CCD camera), CMOS (complementary metal oxide semiconductor) based imaging device (e.g., CMOS camera), photodiode (e.g., photomultiplier tube), electron microscopy, field-effect transistor (e.g., DNA field-effect transistor), ISFET ion sensor (e.g., CHEMFET sensor), or combinations thereof.
[0154] Any suitable MPS method, system, or technical platform for carrying out the methods described herein may be used to obtain nucleic acid sequence reads. Non-exclusive examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and Ion semiconductor-based sequencing (e.g., those developed by Life Technologies), technologies based on WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer (e.g., those developed and marketed by Life Technologies, U.S. Patent Application Publication No. 2013 / 0012399); Polony sequencing, pyrosequencing, Massively Parallelized Signature Sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemosensitive field-effect transistor (CHEMFET) arrays, and electron microscopy-based sequencing (e.g., ZS Examples include Genetics (developed by Halcyon Molecular), nanoball sequencing, or combinations thereof. Other sequencing methods that may be used to carry out the methods described herein include digital PCR, hybridization sequencing, nanopore sequencing, and chromosome-specific sequencing (e.g., using DANSR (Digital Analysis of Selected Regions) techniques).
[0155] In some embodiments, sequence reads are produced, obtained, collected, assembled, manipulated, transformed, processed, and / or provided by the sequence module. Instruments comprising the sequence module may be suitable instruments and / or apparatus for sequencing nucleic acids using sequencing techniques known in the art. In some embodiments, the sequence module may be alignable, assembleable, fragmentable, complementable, reversed It can (complement) and / or error-check (for example, it can error-correct an array read).
[0156] Lead mapping Sequence reads can be mapped, and the number of reads that map to a specific nucleic acid region (e.g., a chromosome or a part thereof) is referred to as the count. Any preferred mapping method (e.g., a process, algorithm, program, software, module, etc., or a combination thereof) can be used. Certain aspects of the mapping process are described hereafter in this specification.
[0157] The mapping of nucleotide sequence reads (i.e., sequence information from fragments whose physical genomic location is unknown) can be performed in several ways, often involving aligning the resulting sequence reads with matching sequences in a reference genome. In such alignment, the sequence reads are typically aligned to a reference sequence, and the sequence reads being aligned are referred to as "mapped," "mapped sequence reads," or "mapped reads." In certain embodiments, mapped sequence reads are referred to as "hits" or "counts." In some embodiments, mapped sequence reads are grouped together according to various parameters and assigned to specific genomic regions, which are discussed in more detail below.
[0158] The terms “aligned,” “alignment,” or “to align” generally refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identity) or a partial match. Alignment can be performed manually or by computer (e.g., software, program, module, or algorithm), a non-exclusive example of which is the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alignment of sequence reads can be a 100% sequence match. In some cases, the alignment is a less than 100% sequence match (i.e., an incomplete match, partial match, or partial alignment). In some embodiments, the alignments are approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some embodiments, the alignments include mismatches. In some embodiments, the alignments include one, two, three, four, or five mismatches. Two or more sequences may be aligned using either strand (e.g., a sense strand or an antisense strand). In certain embodiments, one nucleic acid sequence is aligned with the reverse complementary strand of another nucleic acid sequence.
[0159] Various computer methods can be used to map each sequence read to a specific region. Non-exclusive examples of computer algorithms that can be used to align sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, PROBEMATCH, SOAP, BWA, or SEQMAP, or variations thereof or combinations thereof. In some embodiments, sequence reads can be aligned to sequences in a reference genome. In some embodiments, sequence reads can be found in and / or aligned to sequences in nucleic acid databases known in the art, such as GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (DNA Databank of Japan). BLAST or a similar tool can be used to search for identified sequences against the sequence database. The search hits can then be used, for example, to sort the identified sequences into appropriate regions (described later in this specification).
[0160] In some embodiments, reads may map uniquely or non-uniquely to portions within the reference genome. A read is considered "uniquely mapped" if it aligns with a single sequence in the reference genome. A read is considered "non-uniquely mapped" if it aligns with two or more sequences in the reference genome. In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a certain small mismatch (0-1) may be tolerated to account for single nucleotide polymorphisms that may exist between the reference genome and the reads derived from the individual samples being mapped. In some embodiments, even a small degree of mismatch is not tolerated for reads that map to a reference sequence.
[0161] As used herein, the term “reference genome” may mean any specific known, sequenced, or characterized genome of any organism or virus, whether partial or complete, that can be used as a reference for a specified sequence derived from a subject. For example, reference genomes used for human subjects and many other organisms may be found at the National Center for Biotechnology Information at the World Wide Web URL ncbi.nlm.nih.gov. “Genome” means the complete genetic information of an organism or virus, expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, the reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, the reference genome includes sequences assigned to chromosomes.
[0162] In certain embodiments, mapping ability is evaluated for genomic regions (e.g., parts, genomic segments). Mapping ability is the ability to clearly align nucleotide sequence reads to a portion of the reference genome, typically with a specified number of mismatches (e.g., including 0, 1, 2, or more mismatches). For a given genomic region, the expected mapping ability can be estimated by using a sliding-window approach with a predetermined read length and averaging the mapping ability values obtained at the read levels. Genomic regions containing consecutive unique nucleotide sequences may have high mapping ability values.
[0163] In paired-end sequencing, reads can be mapped to a reference genome by using a suitable mapping and / or alignment program, non-limiting examples of such programs include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754-60), Novoalign Examples include Novocraft (2010), Bowtie (Langmead B et al. (2009) Genome Biol. 10: R25), SOAP2 (Li R et al. (2009) Bioinformatics 25, 1966-67), BFAST (Homer N et al. (2009) PLoS ONE 4, e7767), GASSST (Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534-2540), and MPscan (Rivals E et al. (2009) Lecture Notes in Computer Science 5724, 246-260). Paired-end reads can be mapped and / or aligned using a suitable short-read alignment program. Non-exclusive examples of short-read alignment programs include BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, BWA, CASHX, CUDA-EC, CUSHAW, CUSHAW2, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP, and Geneious. Examples include Assembler, iSAAC, LAST, MAQ, mrFAST, mrsFAST, MOSAIK, MPscan, Novoalign, NovoalignCS, Novocraft, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3, SOCS, SSAHA, SSAHA2, Stampy, SToRM, Subread, Subjunc, Taipan, UGENE, VelociMapper, TimeLogic, XpressAlign, ZOOM, and combinations thereof. Paired-end reads are often mapped to opposite ends of the same polynucleotide fragment according to the reference genome. In some embodiments, readmates are mapped independently. In some embodiments, information from both sequence reads (i.e., from each end) is considered in the mapping process.A reference genome is often used to determine and / or predict the sequences of nucleic acids located between paired-end readmates. The term “mismatched read pair,” as used herein, refers to a paired-end read containing a pair of readmates where one or both readmates do not clearly map to the same region of the reference genome, partially defined by a segment of consecutive nucleotides. In some embodiments, a mismatched read pair is a paired-end readmate that maps to an unexpected location in the reference genome. Non-limiting examples of an unexpected location in the reference genome include (i) two different chromosomes, (ii) locations beyond a given fragment size (e.g., beyond 300 bp, beyond 500 bp, beyond 1000 bp, beyond 5000 bp, or beyond 10,000 bp), (iii) orientations that do not match the reference sequence (e.g., reverse orientation), or combinations thereof. In some embodiments, a mismatched readmate is identified according to the length (e.g., average length, given fragment size) or expected length of a template polynucleotide fragment in the sample. For example, readmates that map to positions farther than the average or expected length of polynucleotide fragments in a sample may be identified as mismatched read pairs. Read pairs that map in opposite directions may be determined by obtaining the reverse complement of one of those reads and comparing the alignment of both reads using the same strand of the reference sequence. Mismatched read pairs may be identified by any preferred method and / or algorithm known in the art or described herein (e.g., SVDetect, Lumpy, BreakDancer, BreakDancerMax, CREST, DELLY, etc., or a combination thereof).
[0164] portion In some embodiments, mapped sequence reads are grouped together according to various parameters and assigned to a specific genomic region (e.g., a reference genome region). The “region” may also be referred to herein as a “genomic section,” “bin,” “partition,” “reference genome region,” “chromosome region,” or “genomic region.”
[0165] A segment is often defined by dividing the genome according to one or more features. Non-limiting examples of certain features of a segment include length (e.g., default length, non-default length) and other structural features. A genomic segment may include one or more of the following features: default length, non-default length, random length, non-random length, equal length, unequal length (e.g., at least two segments of the genomic segment are of unequal length), non-overlapping (e.g., the 3' end of one segment of the genomic segment may be adjacent to the 5' end of an adjacent segment), overlapping (e.g., at least two segments of the genomic segment overlap), contiguous, continuous, non-contiguous, and non-contiguous. The genomic portion can be approximately 1 to 1,000 kilobase lengths (for example, approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900 kilobase lengths), approximately 5 to 500 kilobase lengths, approximately 10 to 100 kilobase lengths, or approximately 40 to 60 kilobase lengths.
[0166] Splitting may or may not be based on features relating to certain information (e.g., information content and information growth). Non-limiting examples of features relating to certain information include alignment speed and / or convenience, sequencing coverage variability, GC content (e.g., stratified GC content, specific GC content, high or low GC content), GC content uniformity, other measures of sequence content (e.g., ratio of individual nucleotides, ratio of pyrimidines or purines, ratio of native to non-native nucleic acids, ratio of methylated nucleotides and CpG content), methylation status, double-strand melting temperature, amenability for sequencing or PCR, uncertainty assigned to individual parts of the reference genome, and / or targeted search for specific features. In some embodiments, information content may be quantified using p-value profiles that measure the significance of specific genomic locations to distinguish between groups of confirmed normal and abnormal subjects (e.g., euploid subjects and trisomic subjects, respectively).
[0167] In some embodiments, splitting the genome can eliminate similar regions (e.g., identical or homologous regions or sequences) across the genome, leaving only unique regions. The regions removed in splitting may be located within a single chromosome, one or more chromosomes, or span multiple chromosomes. In some embodiments, the split genome is often reduced and optimized for faster alignment to focus on uniquely identifiable sequences.
[0168] In some embodiments, genomic portions arise from divisions based on a predetermined non-overlapping size, thereby resulting in a contiguous series of non-overlapping portions of a predetermined length. Such portions are often shorter than chromosomes and often shorter than regions of copy number variation (or copy number change) (e.g., duplicated or deleted regions), the latter of which may be referred to as segments. A “segment” or “genomic segment” often contains two or more genomic portions of a predetermined length, and often contains two or more contiguous portions of a predetermined length (e.g., about 2 to about 100 such portions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 such portions)).
[0169] Multiple parts may be analyzed in groups, and reads mapped to parts may be quantified according to specific groups of genomic parts. If parts are divided by structural features and correspond to regions in the genome, parts may be grouped into one or more segments and / or one or more regions. Non-limiting examples of regions include subchromosomes (i.e., shorter than chromosomes), chromosomes, autosomes, sex chromosomes, and combinations thereof. One or more subchromosomal regions may be genes, gene fragments, regulatory sequences, introns, exons, segments (e.g., segments spanning copy number variation regions; segments spanning copy number variation regions), microduplications, microdeletions, etc. Regions may be smaller than or the same size as the chromosome of interest, or smaller than or the same size as the reference chromosome.
[0170] Filtering and / or selecting parts In some embodiments, one or more processing steps may include one or more partial filtering steps and / or partial selection steps. The term “filtering,” as used herein, means removing a portion or part of a reference genome from consideration. In certain embodiments, one or more portions are filtered (e.g., subjected to a filtering process) to provide filtered portions. In some embodiments, a filtering process removes a particular portion and retains a portion (e.g., a subset of a portion). After a filtering process, the retained portion is often referred to herein as the filtered portion.
[0171] Parts of the reference genome may be selected for removal based on any preferred criteria, including, but not limited to, redundant data (e.g., redundant or overlapping mapped reads), uninformative data (e.g., parts of the reference genome with a median count of zero), parts of the reference genome containing over- or under-presented sequences, noisy data, or a combination of the foregoing. The filtering process often involves removing one or more parts of the reference genome from consideration and subtracting the counts of one or more parts of the reference genome selected for removal from the counted or totaled counts of the reference genome, chromosome, or parts of the genome under consideration. In some embodiments, parts of the reference genome may be removed sequentially (e.g., one by one to allow evaluation of the impact of removing each individual part), and in certain embodiments, all parts of the reference genome marked for removal may be removed simultaneously. In some embodiments, parts of the reference genome characterized by variance above or below a certain level are removed, which is sometimes referred to herein as filtering of “noisy” parts of the reference genome. In certain embodiments, the filtering process includes obtaining from the dataset data points that deviate from the mean profile level of a given set of profiles for each set of profiles, and in certain embodiments, the filtering process includes removing from the dataset data points that do not deviate from the mean profile level of a given set of profiles for each set of profiles. In some embodiments, the filtering process is used to reduce the number of candidate regions of the reference genome that are analyzed for the presence or absence of gene mutations / genetic alterations and / or copy number alterations (e.g., aneuploidy, microdeletions, microduplications).Reducing the number of candidate reference genome segments analyzed for the presence or absence of gene mutations / mutations and / or copy number variations often lowers the complexity and / or dimensionality of the dataset, sometimes increasing the speed of searching for and / or identifying gene mutations / mutations and / or copy number variations by two orders of magnitude or more.
[0172] The portions may be processed (e.g., filtered and / or selected) by any preferred method and according to any preferred parameters. Non-limiting examples of features and / or parameters that may be used to filter and / or select portions include redundant data (e.g., redundant or overlapping mapped reads), uninformative data (e.g., portions of the reference genome with zero mapped counts), portions of the reference genome containing over- or under-presented sequences, noisy data, counts, count variability, coverage, mapping, variability, repeatability, read density, read density variability, level of uncertainty, guanine-cytosine (GC) content, CCF fragment length and / or read length (e.g., fragment length ratio (FLR), fetal ratio statistic (FRS)), DNase I sensitivity, methylation status, acetylation, histone distribution, chromatin structure, repeat percentage, etc., or combinations thereof. The portions may be filtered and / or selected according to any preferred features or parameters that correlate with the features or parameters enumerated or described herein. Parts may be filtered and / or selected according to features or parameters specific to a part (e.g., when measured on a single part relating to multiple samples) and / or features or parameters specific to a sample (e.g., when measured on multiple parts within a single sample). In some embodiments, parts are filtered and / or removed according to relatively low mapping ability, relatively large variability, high levels of uncertainty, relatively long CCF fragment lengths (e.g., low FRS, low FLR), relatively high proportion of repeating sequences, high GC content, low GC content, low count, zero count, high count, etc., or a combination thereof. In some embodiments, parts (e.g., subsets of parts) are selected according to a preferred level of mapping ability, variability, level of uncertainty, proportion of repeating sequences, count, GC content, etc., or a combination thereof.In some embodiments, parts (e.g., subsets of parts) are selected according to relatively short CCF fragment lengths (e.g., high FRS, high FLR). Counts and / or reads mapped to parts may be processed (e.g., normalized) before and / or after filtering or selecting parts (e.g., subsets of parts). In some embodiments, counts and / or reads mapped to parts are not processed before and / or after filtering or selecting parts (e.g., subsets of parts).
[0173] In some embodiments, portions may be filtered according to a measure of error (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation and / or mean absolute deviation (MAD). In certain cases, the measure of error may refer to count variability. In some embodiments, portions are filtered according to count variability. In certain embodiments, count variability is a measure of error determined for counts mapped to portions of the reference genome (i.e., portions) for multiple samples (e.g., multiple subjects, e.g., multiple samples obtained from 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more subjects). In some embodiments, In some embodiments, portions of the count variation above a predetermined upper range are filtered out (e.g., excluded from consideration). In some embodiments, portions of the count variation below a predetermined lower range are filtered out (e.g., excluded from consideration). In some embodiments, portions of the count variation outside a predetermined range are filtered out (e.g., excluded from consideration). In some embodiments, portions of the count variation within a predetermined range are selected (e.g., used to determine the presence or absence of copy number variation). In some embodiments, the count variation of a portion exhibits a distribution (e.g., a normal distribution). In some embodiments, a portion within a certain quantile of that distribution is selected. In some embodiments, a portion within the 99th percentile of the distribution of count variation is selected.
[0174] In some embodiments, a method for classifying the presence or absence of copy number variations in a subchromosomal region of a test sample includes a step of identification using a segmentation process. In some embodiments, the presence or absence of a copy number variation segment may be in a region containing a first set of genomic segments, which includes at least a portion of the subchromosomal region of interest. As an illustrative example, the region containing the first set of genomic segments is the region enclosed by the black dashed line in Figure 4. In some embodiments, the first set of genomic segments is a portion within a region of a chromosome in which copy number variations associated with the phenotype of interest are expected to exist. In some embodiments, such genomic segments can often be obtained by mining public disease databases, such as the International Standards of Cytogenomic Arrays database (ISCA). In some embodiments, the genomic segments used herein may be identified within the subchromosomal region of interest by a circular binary segmentation (CBS) algorithm. In one embodiment, the phenotype is a microdeletion syndrome. In one embodiment, the first set of genomic segments is one or more genomic segments selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23.
[0175] In some embodiments, a method for classifying the presence or absence of copy number variation in a subchromosomal region of a test sample includes the step of providing quantitative values of sequence reads for a subregion within the subchromosomal region, which comprises a set of genomic subregions. The genomic subregion includes a portion of the reference genome to which the obtained sequence reads are mapped to nucleic acids in the test sample. In some embodiments, the set is a predetermined set of genomic subregions. As an illustrative example, the subchromosomal region is the area enclosed by the black dashed line in Figure 4.
[0176] In some embodiments, a given set of genomic regions is identified according to one or more precision measures for multiple samples in a training set, each of which is classified as having copy number variation in the subchromosomal region of interest. Precision measures may include, but are not limited to, sensitivity, specificity, standard deviation, median absolute deviation (MAD), determinism, confidence, uncertainty, coefficient of variation (CV), confidence level, confidence interval (e.g., about 95% confidence interval), standard score (e.g., z score), chi score, phi score, t-test result, p-value, ploidy, fitted minority ratio, area ratio, median level, etc., or combinations thereof, as described in detail herein. In some embodiments, precision measures include sensitivity. The genome portion is selected based on the fact that it provides an optimal precision measure, i.e., a precision measure equal to or higher than a given threshold (which is considered the minimum requirement for detecting the presence or absence of copy number variation with reasonable precision). For example, when sensitivity is used as the precision measure, the threshold can be any number between 70% and 100%, e.g., 75% and 99%, 80% and 98%, or 85% and 95%.
[0177] In one embodiment, a given set of genome portions is identified by a process comprising: 1) providing a plurality of candidate subregions within a subchromosomal region; 2) providing one or more precision measures for each of the plurality of candidate subregions for a plurality of samples in a training set, wherein each of the plurality of samples is classified as having copy number variation in the subchromosomal region; and 3) identifying the subregion in (a) as the subregion that provides optimal precision according to one or more precision measures.
[0178] Sequence reads derived from any suitable number of samples may be used to identify a subset of portions that satisfy one or more criteria, parameters, and / or features described herein. Sequence reads from a group of samples derived from multiple subjects may be used. In some embodiments, the multiple subjects include pregnant women. In some embodiments, the multiple subjects include healthy subjects. In some embodiments, the multiple subjects include cancer patients. One or more samples derived from each of multiple subjects (e.g., 1 to about 20 samples derived from each subject (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 samples)) may be dealt with, and a suitable number of subjects (e.g., about 2 to about 10,000 subjects (e.g., about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 subjects)) may be dealt with. In some embodiments, sequence reads from the same test sample derived from the same subject are mapped to a portion in the reference genome and used to generate a subset of that portion.
[0179] The portions may be selected and / or filtered by any preferred method. In some embodiments, portions are selected according to a visual inspection of data, graphs, plots and / or charts. In certain embodiments, portions are selected and / or filtered (e.g., partially) by a system or device having one or more microprocessors and memory. In some embodiments, portions are selected and / or filtered (e.g., partially) by a non-temporary computer-readable storage medium on which an executable program is stored, instructing the microprocessor to perform the selection and / or filtering.
[0180] In some embodiments, sequence reads derived from a sample are mapped to all or most of a reference genome, and then a pre-selected subset of that subset is chosen. For example, a subset of a subset of a subset of a subset of a fragment at a certain length threshold may be chosen, where reads from the fragment preferentially map to that subset. A particular method for pre-selecting a subset of a subset of a subset is described in U.S. Patent Application Publication No. 2014 / 0180594 (incorporated herein by reference). Reads from the selected subset of a
[0181] In some embodiments, portions related to read density (e.g., where the read density is relative to a portion) are removed by the filtering process, and the read densities related to the removed portions are not included in the determination of the presence or absence of copy number changes (e.g., chromosomal aneuploidy, microduplication, microdeletion). In some embodiments, the read density profile includes and / or consists of the read densities of the filtered portions. Portions may be filtered according to the distribution of counts and / or read densities. In some embodiments, portions are filtered according to the distribution of counts and / or read densities, where those counts and / or read densities are obtained from one or more reference samples. One or more reference samples may be referred to herein as a training set. In some embodiments, portions are filtered according to the distribution of counts and / or read densities, where those counts and / or read densities are obtained from one or more test samples. In some embodiments, portions are filtered according to a measure of uncertainty for the read density distribution. In certain embodiments, portions showing large deviations in read density are removed by the filtering process. For example, a distribution of read densities (e.g., the mean or median distribution of the mean read densities) can be determined, where each read density in that distribution maps to the same region. A measure of uncertainty (e.g., MAD) can be determined by comparing the read density distributions for multiple samples, where each region of the genome is associated with a measure of uncertainty. As in the example above, regions can be filtered according to a measure of uncertainty (e.g., standard deviation (SD), MAD) associated with each region and a predetermined threshold. In a particular case, regions containing MAD values within an acceptable range are retained, and regions containing MAD values outside an acceptable range are removed from consideration by the filtering process.In some embodiments, as in the examples described above, portions containing read density values outside a given uncertainty scale (e.g., median, mean, or average read density) are often removed from consideration by the filtering process. In some embodiments, portions containing read density values outside the interquartile range of a distribution (e.g., median, mean, or average read density) are removed from consideration by the filtering process. In some embodiments, portions containing read density values more than 2, 3, 4, or 5 times outside the interquartile range of a distribution are removed from consideration by the filtering process. In some embodiments, portions containing read density values more than 2 sigma, 3 sigma, 4 sigma, 5 sigma, 6 sigma, 7 sigma, or 8 sigma (e.g., sigma is the range defined by the standard deviation) are removed from consideration by the filtering process.
[0182] Quantitative values of sequence reads Sequence reads mapped or segmented based on selected features or variables can, in some embodiments, be quantified to measure the amount or number of reads mapped to one or more parts (e.g., parts of a reference genome). In certain embodiments, the amount of sequence reads mapped to a particular part or segment is referred to as the count or read density.
[0183] Counts are often associated with genomic regions. In some embodiments, counts are measured from some or all of the sequence reads mapped to (i.e., associated with) a region. In certain embodiments, counts are measured from some or all of the sequence reads mapped to a group of regions (e.g., a portion within a segment or region (as described herein)).
[0184] The count may be measured by a preferred method, calculation, or mathematical process. The count may be the direct sum of all sequence reads mapped to a genomic portion or group of genomic portions corresponding to a segment, a group of portions corresponding to a subregion of the genome (e.g., copy number variant regions, copy number change regions, copy number duplicate regions, copy number deletion regions, microduplication regions, microdeletion regions, chromosomal regions, autosomal regions, sex chromosome regions), and / or a group of portions corresponding to the genome. The quantitative value of a read may be a ratio, which may be the ratio of the quantitative value for a portion in region a to the quantitative value for a portion in region b. Region a may be a single portion, a segment region, a copy number variant region, a copy number change region, a copy number duplicate region, a copy number deletion region, a microduplication region, microdeletion region, chromosomal regions, autosomal regions, and / or sex chromosome regions. Region b may independently be a single part, a segment region, a copy number variation region, a copy number change region, a copy number duplication region, a copy number deletion region, a microduplication region, a microdeletion region, a chromosome region, an autosomal region, a sex chromosome region, a region containing all autosomes, a region containing sex chromosomes, and / or a region containing all chromosomes.
[0185] In some embodiments, the count is obtained from raw and / or filtered sequence reads. In certain embodiments, the count is the mean, average, or sum of sequence reads mapped to a genomic portion or group of genomic portions (e.g., a genomic portion within a region). In some embodiments, the count is related to indeterminate values. The count may be adjusted. The count may be adjusted according to sequence reads associated with a genomic portion or group of portions that have been weighted, removed, filtered, normalized, adjusted, averaged, derived as mean, derived as median, added, or a combination thereof.
[0186] The quantitative value of sequence reads may sometimes be read density. Read density can be measured and / or generated for one or more segments of the genome. In certain cases, read density can be measured and / or generated for one or more chromosomes. In some embodiments, read density includes a quantitative measure of the count of sequence reads mapped to a segment or portion of a reference genome. Read density can be measured by a preferred process. In some embodiments, read density is measured by a preferred distribution and / or a preferred distribution function. Non-restrictive examples of distribution functions include any preferred distribution or combination thereof, such as a probability function, probability distribution function, probability density function (PDF), kernel density function (kernel density estimate), cumulative distribution function, probability mass function, discrete probability distribution, absolute continuous univariate distribution, etc. Read density may be a density estimate derived from a preferred probability density function. The density estimate is the construction of an estimate of the latent probability density function based on observed data. In some embodiments, read density includes density estimates (e.g., probability density estimate, kernel density estimate). Read density may be generated by a process that includes the step of generating a density estimate for each of one or more parts of the genome (where each part includes a count of sequence reads). Read density may be generated for normalized and / or weighted counts mapped to the parts or segments. In some cases, each read mapped to a part or segment may contribute to a read density that is equal to its weight (e.g., count) obtained from the normalization process described herein. In some embodiments, the read density for one or more parts or segments is adjusted. Read density may be adjusted by preferred methods. For example, the read density for one or more parts may be weighted and / or normalized.
[0187] Reads quantified for a given portion or segment may originate from one or different sources. In one example, reads may be obtained from nucleic acids derived from a subject who has or is suspected of having cancer. In such a situation, reads mapped to one or more portions are often representative of both healthy cells (i.e., non-cancerous cells) and cancer cells (e.g., tumor cells). In a particular embodiment, some of the reads mapped to a portion may originate from cancer cell nucleic acids, and some of the reads mapped to the same portion may originate from non-cancerous cell nucleic acids. In another example, reads may be obtained from nucleic acid samples derived from a pregnant woman with a fetus. In such a situation, reads mapped to one or more portions may often represent both the fetus and the mother of the fetus (e.g., the pregnant subject). In a particular embodiment, some of the reads mapped to a portion may originate from the fetal genome, and some of the reads mapped to the same portion may originate from the maternal genome.
[0188] level In some embodiments, a value (e.g., a number, a quantitative value) is attributed to a level. The level can be determined by a preferred method, operation, or mathematical process (e.g., a processed level). Often, the level is a count for a subset (e.g., a normalized count) or is derived from such a count. In some embodiments, the level of a subset is substantially equal to the total number of counts mapped to the subset (e.g., counts, normalized counts). The level is often determined from counts that have been processed, transformed, or manipulated by a preferred method, operation, or mathematical process known in the art. In some embodiments, a level is derived from a processed count, and non-limiting examples of processed counts include weighted counts, removed counts, filtered counts, normalized counts, adjusted counts, averaged counts, counts derived as mean values (e.g., mean level), added counts, subtracted counts, transformed counts, or combinations thereof. In some embodiments, a level includes a normalized count (e.g., a normalized count of a subset). Some levels may be for counts normalized by a preferred process, non-limiting examples of such processes are described herein. Some levels may include normalized counts or relative amounts of counts. In some embodiments, some levels are for two or more averaged portions of counts or normalized counts, and these levels are referred to as mean levels. In some embodiments, some levels are for a subset having a mean of the counts or a mean of the normalized counts, and are referred to as mean levels. In some embodiments, some levels are derived for portions containing raw counts and / or filtered counts. In some embodiments, some levels are based on raw counts. In some embodiments, some levels relate to indeterminate values (e.g., standard deviation, MAD). In some embodiments, some levels are represented by Z-scores or p-values.
[0189] A level for one or more parts is synonymous with “genome section level” in this specification. The term “level” may be synonymous with the term “height” when used herein. The meaning of the term “level” can be determined from the context in which it is used. For example, when the term “level” is used in the context of parts, profiles, reads and / or counts, it often means height. When the term “level” is used in the context of substances or compositions (e.g., RNA level, plexing level, quantity), it often means quantity. When the term “level” is used in the context of uncertainty (e.g., error level, confidence level, deviation level, uncertainty level), it often means quantity.
[0190] Normalized or unnormalized counts for two or more levels (e.g., two or more levels in a profile) can sometimes be manipulated mathematically according to the levels (e.g., they can be added, multiplied, averaged, normalized, etc., or a combination thereof). For example, normalized or unnormalized counts for two or more levels can be normalized according to one, some, or all of the levels in a profile. In some embodiments, normalized or unnormalized counts for all levels in a profile are normalized according to one level in that profile. In some embodiments, normalized or unnormalized counts for the first (fist) level in a profile are normalized according to the normalized or unnormalized counts for the second level in that profile.
[0191] Non-limiting examples of levels (e.g., first level, second level) include levels for a subset containing processed counts, levels for a subset containing the mean, median, or average of the counts, levels for a subset containing normalized counts, or any combination thereof. In some embodiments, the first and second levels in a profile are derived from counts of portions mapped to the same chromosome. In some embodiments, the first and second levels in a profile are derived from counts of portions mapped to different chromosomes.
[0192] In some embodiments, a level is determined from normalized or unnormalized counts mapped to one or more parts. In some embodiments, a level is determined from normalized or unnormalized counts mapped to two or more parts, where the normalized counts for each part are often approximately the same. Variation in counts (e.g., normalized counts) may exist in a subset of a level. In a subset of a level, there may be one or more parts that have counts significantly different from the other parts of the set (e.g., peaks and / or dips). Any number of normalized or unnormalized counts associated with any number of parts may define a level.
[0193] In some embodiments, one or more levels may be determined from all or some normalized or unnormalized counts of a portion of a genome. Often, levels may be determined from all or some normalized or unnormalized counts of a chromosome or a portion of it. In some embodiments, two or more counts derived from two or more portions (e.g., a set of portions) determine a level. In some embodiments, two or more counts (e.g., counts from two or more portions) determine a level. In some embodiments, counts from 2 to about 100,000 portions determine a level. In some embodiments, counts from portions of 2 to approximately 50,000, 2 to approximately 40,000, 2 to approximately 30,000, 2 to approximately 20,000, 2 to approximately 10,000, 2 to approximately 5,000, 2 to approximately 2,500, 2 to approximately 1,250, 2 to approximately 1,000, 2 to approximately 500, 2 to approximately 250, 2 to approximately 100, or 2 to approximately 60 determine the level. In some embodiments, counts from portions of approximately 10 to approximately 50 determine the level. In some embodiments, counts from portions of approximately 20 to approximately 40 or more determine the level. In some embodiments, a certain level includes counts from approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more parts. In some embodiments, a certain level corresponds to a subset (e.g., a subset of a reference genome, a subset of a chromosome, or a subset of a part of a chromosome).
[0194] In some embodiments, a level is determined for normalized or unnormalized counts of contiguous portions. In some embodiments, contiguous portions (e.g., a set of portions) correspond to adjacent regions of a genome or adjacent regions of a chromosome or gene. For example, two or more contiguous portions may be a sequence assembly of DNA sequences longer than each portion when aligned by merging those portions end-to-end. For example, two or more contiguous portions may be an intact genome, chromosome, gene, intron, exon, or a part thereof. In some embodiments, a level is determined from a collection (e.g., a set) of contiguous and / or non-contiguous portions.
[0195] Data processing and normalization Counted mapped sequence reads are referred to herein as raw data, because they correspond to unmanipulated counts (e.g., raw counts). In some embodiments, sequence read data in a dataset may be further processed (e.g., mathematically and / or statistically manipulated) and / or presented to facilitate the presentation of outcomes. In certain embodiments, datasets containing larger datasets may benefit from preprocessing to facilitate further analysis. Preprocessing of a dataset may include removing redundant and / or non-informational portions or portions of the reference genome (e.g., portions of the reference genome with non-informational data, redundant mapped reads, portions with a median count of zero, over-presented or under-presented sequences). Not limited to theory, data processing and / or preprocessing may (i) remove noisy data, (ii) remove non-informational data, (iii) remove redundant data, (iv) reduce the complexity of larger datasets, and / or (v) facilitate the transformation of data from one form to one or more other forms. The terms “preprocessing” and “processing” are collectively referred to as “processing” in this specification when used in relation to data or datasets. Processing can make data applicable to further analysis and, in some embodiments, can produce outcomes. In some embodiments, one or more processing methods or all of the processing methods (e.g., normalization methods, partial filtering, mapping, validation, etc., or combinations thereof) are performed by a memory-connected processor, microprocessor, computer, and / or device controlled by a microprocessor.
[0196] The term “noisy data,” as used herein, refers to (a) data with significant variance between data points when analyzed or plotted, (b) data with significant standard deviation (e.g., a standard deviation greater than 3), (c) data with significant standard error of the mean, and combinations thereof. Noisy data may arise due to the quantity and / or quality of the starting material (e.g., nucleic acid sample), or as part of the process for preparing or replicating DNA used to generate sequence reads. In certain embodiments, noise may result from certain sequences that are overpresented when prepared using PCR-based methods. The methods described herein may reduce or eliminate the involvement of noisy data and thus reduce its impact on the outcomes provided.
[0197] The terms “data of no informational value,” “reference genome portion of no informational value,” and “portion of no informational value,” as used herein, refer to a portion or data derived therefrom that has a numerical value that is significantly different from a given threshold or does not fall within a given cutoff range of values. The terms “threshold” and “threshold” as used herein refer to any numerical value calculated using a qualifying dataset that serves as a limit for diagnosing a gene variant or gene change (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormality, etc.). In certain embodiments, the threshold is exceeded by the results obtained by the methods described herein, and the subject is diagnosed with a copy number variation. Thresholds or ranges of values are often calculated in some embodiments by mathematically and / or statistically manipulating sequence read data (e.g., sequence read data from a reference and / or subject), and in certain embodiments, the sequence read data manipulated to generate thresholds or ranges of values is sequence read data (e.g., sequence read data from a reference and / or subject). In some embodiments, an indeterminate value is determined. An uncertainty value is generally a measure of variance or error, and can be any suitable measure of variance or error. In some embodiments, the uncertainty value is the standard deviation, standard error, calculated variance, p-value, or mean absolute deviation (MAD). In some embodiments, the uncertainty value may be calculated according to the formulas described herein.
[0198] Any suitable procedure may be used to process the datasets described herein. Non-limiting examples of suitable procedures for processing the datasets include filtering, normalization, weighting, monitoring of peak height, monitoring of peak area, monitoring of peak edge, peak level analysis, peak width analysis, peak edge position analysis, peak lateral tolerances, measurement of area ratios, mathematical processing of the data, statistical processing of the data, application of statistical algorithms, analysis with fixed variables, analysis with optimized variables, plotting of the data to identify patterns or trends for further processing, and combinations thereof. In some embodiments, the datasets are processed based on various features (e.g., GC content, mapped redundant reads, centromere regions, telomere regions, etc., and combinations thereof) and / or variables (e.g., sex of the subject, age of the subject, ploidy of the subject, percentage contribution of cancer cell nucleic acids, sex of the fetus, age of the mother, ploidy of the mother, percentage contribution of fetal nucleic acids, etc., or combinations thereof). In certain embodiments, processing datasets as described herein can reduce the complexity and / or dimensionality of large and / or complex datasets. Non-limiting examples of complex datasets include sequence read data generated from one or more test subjects and multiple reference subjects of different ages and ethnic backgrounds. In some embodiments, a dataset may contain thousands to millions of sequence reads for each test subject and / or each reference subject.
[0199] Data processing can be performed in any number of steps in a particular embodiment. For example, in some embodiments, data may be processed using only one processing step, and in some particular embodiments, data may be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some embodiments, the processing step may be the same step repeated two or more times (e.g., filtering two or more times, normalization two or more times), and in certain embodiments, the processing step may be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, normalization; normalization of peak height and peak edge, monitoring; filtering, normalization, normalization relative to a reference, statistical operations to determine p-values, etc.). In some embodiments, any preferred number and / or combination of the same or different processing steps may be used to process the sequence read data to facilitate the provision of outcomes. In certain embodiments, processing the dataset according to the criteria described herein may reduce the complexity and / or dimensionality of the dataset.
[0200] In some embodiments, one or more processing steps may include one or more normalization steps. Normalization may be performed by preferred methods described herein or known in the art. In certain embodiments, normalization includes adjusting values measured on different scales to a conceptually common scale. In certain embodiments, normalization includes advanced mathematical adjustments to bring the probability distribution of the adjusted values into alignment. In some embodiments, normalization includes aligning the distributions to a normal distribution. In certain embodiments, normalization includes mathematical adjustments to enable comparison of corresponding normalized values to different datasets in order to eliminate the effect of certain overall influences (e.g., errors and exceptions). In certain embodiments, normalization includes scaling. Normalization may sometimes include division of one or more datasets by a given variable or expression. Normalization may sometimes include subtraction of one or more datasets by a given variable or expression. Non-restrictive examples of normalization methods include partwise normalization, GC content normalization, median count normalization (median bin count, median subcount), linear and nonlinear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatter plot smoothing), principal component normalization, repeat mask (RM), GC normalization and repeat mask (GCRM), cQn, and / or combinations thereof. In some embodiments, the presence or absence of copy number variations (e.g., aneuploidy, minute overlaps, minute deletions) is determined using normalization methods (e.g., partwise normalization, GC content normalization, median count normalization (median bin count, median part count), linear and nonlinear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatter plot smoothing), principal component normalization, repeat mask (RM), GC normalization and repeat mask (GCRM), cQn, normalization methods known in the art, and / or combinations thereof). Certain examples of normalization processes that may be used, such as LOESS normalization, principal component normalization, and hybrid normalization methods, will be described in more detail later in this specification.Certain aspects of the normalization process are also described, for example, in international patent application publication numbers WO2013 / 052913 and WO2015 / 051163 (each incorporated herein by reference).
[0201] An arbitrary number of preferred normalizations can be used. In some embodiments, a dataset may be normalized only once or more times, five or more times, ten or more times, or even twenty or more times. A dataset may be normalized to a value (e.g., a normalized value) that represents any preferred feature or variable (e.g., sample data, reference data, or both). Non-limiting examples of types of data normalization that may be used include: normalizing raw count data for one or more selected test or reference portions to the total number of counts mapped to the entire chromosome or genome to which the selected portion or portion is mapped; normalizing raw count data for one or more selected portions to the median of the reference counts for one or more portions or chromosomes to which the selected portion is mapped; normalizing raw count data to pre-normalized data or its derivative; and normalizing pre-normalized data to one or more other predetermined normalization variables. Normalization of a dataset may have the effect of decoupling statistical errors, depending on the feature or characteristic selected as the predetermined normalization variable. Normalizing a dataset can sometimes allow for the comparison of data properties of data with different scales by bringing the data to a common scale (e.g., a given normalization variable). In some embodiments, one or more normalizations of statistically derived values may be used to minimize data differences and reduce the importance of out-of-range data. Normalizing a portion or a portion of a reference genome to normalization values is sometimes referred to as "part-by-part normalization."
[0202] In certain embodiments, the processing steps may include one or more mathematical and / or statistical operations. Any suitable mathematical and / or statistical operation may be used alone or in combination to analyze and / or manipulate the datasets described herein. Any number of suitable mathematical and / or statistical operations may be used. In some embodiments, the dataset may be mathematically and / or statistically manipulated only once or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical operations that may be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimators, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi-values, phi-values, peak level analysis, peak edge position determination, peak area ratio calculation, chromosome-level median analysis, mean absolute deviation calculation, sum of squared residuals, mean, standard deviation, standard error, etc. or combinations thereof. Mathematical and / or statistical operations may be performed on all or part of the sequence read data or on the processed data. Non-exclusive examples of variables or features of a dataset that can be statistically manipulated include raw counts, filtered counts, normalized counts, peak height, peak width, peak area, peak edges, lateral tolerance, p-values, median levels, mean levels, distribution of counts within genomic regions, relative presentation of nucleic acid species, or combinations thereof.
[0203] In some embodiments, the processing steps may include the use of one or more statistical algorithms. Any suitable statistical algorithm may be used alone or in combination to analyze and / or manipulate the datasets described herein. Any number of suitable statistical algorithms may be used. In some embodiments, the dataset may be analyzed using one or more, five or more, ten or more, or twenty or more statistical algorithms. Non-exclusive examples of statistical algorithms suitable for use with the methods described herein include principal component analysis, decision trees, alternative null hypotheses, multiple comparisons, summative tests, the Behrens-Fisher problem, bootstrapping, Fisher's method for combining independent significance tests, null hypotheses, Type I errors, Type II errors, exact tests, one-sample Z-tests, two-sample Z-tests, one-sample t-tests, paired t-tests, pooled two-sample t-tests with equal variances, unpooled two-sample t-tests with unequal variances, one-proportion Z-tests, pooled two-proportion Z-tests, unpooled two-proportion Z-tests, one-sample chi-squared tests, two-sample F-tests for equalizing variances, confidence intervals, credible intervals, significance, meta-analysis, simple linear regression, robust linear regression, etc., or combinations thereof. Non-limiting examples of variables or features of a dataset that can be analyzed using statistical algorithms include raw counts, filtered counts, normalized counts, peak height, peak width, peak edge, lateral tolerance, p-value, median level, mean level, distribution of counts within a genomic region, relative presentation of nucleic acid species, or combinations thereof.
[0204] In certain embodiments, a dataset may be analyzed by using multiple (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, K-nearest neighbors, logistic regression, and / or smoothing methods) and / or mathematical and / or statistical operations (e.g., referred to herein as operations). In some embodiments, the use of multiple operations may generate an N-dimensional space that can be used to provide outcomes. In certain embodiments, analysis of a dataset by using multiple operations may reduce the complexity and / or dimensionality of that dataset. For example, by using multiple operations on a reference dataset, an N-dimensional space (e.g., a probability plot) may be generated that can be used to represent the presence or absence of gene mutations / gene changes and / or copy number changes, depending on the state of the reference sample (e.g., positive or negative for a selected copy number change). Analysis of test samples using substantially similar sets of operations may be used to generate N-dimensional points for each test sample. The complexity and / or dimensionality of the test subject dataset may be reduced to a single value or an N-dimensional point that can be readily compared to the N-dimensional space generated from the reference data. Data from test samples that fall into the N-dimensional space occupied by the reference subject data suggest a genetic state substantially similar to that of the reference subject. Data from test samples that do not fall into the N-dimensional space occupied by the reference subject data suggest a genetic state substantially different from that of the reference subject. In some embodiments, the reference is euploid or does not have any particular genetic mutations / genetic alterations and / or copy number alterations and / or medical conditions.
[0205] After a dataset has been counted, filtered as necessary, normalized, and weighted as necessary, the processed dataset may, in some embodiments, be further manipulated by one or more filtering and / or normalization and / or weighting steps. The dataset further manipulated by one or more filtering and / or normalization and / or weighting steps may, in certain embodiments, be used to generate a profile. One or more filtering and / or normalization and / or weighting steps may, in some embodiments, reduce the complexity and / or dimensionality of the dataset. Outcomes may be provided based on the dataset with reduced complexity and / or dimensionality. In some embodiments, for example, a plot of the profile of the processed data further manipulated by weighting is generated to facilitate classification and / or outcome provision. Outcomes may be provided, for example, based on a plot of the profile of the weighted data.
[0206] Filtering or weighting of segments may be performed at one or more suitable points in the analysis. For example, segments may be filtered or weighted before or after sequence reads are mapped to reference genome segments. In some embodiments, segments may be filtered or weighted before or after experimental biases for individual genome segments are determined. In certain embodiments, segments may be filtered or weighted before or after levels are calculated.
[0207] After the dataset has been counted, filtered as necessary, normalized, and weighted as necessary, the processed dataset may, in some embodiments, be manipulated by one or more mathematical and / or statistical operations (e.g., statistical functions or statistical algorithms). In certain embodiments, the processed dataset may be further manipulated by calculating Z-scores for one or more selected parts, chromosomes, or chromosomal segments. In some embodiments, the processed dataset may be further manipulated by calculating P-values. In certain embodiments, the mathematical and / or statistical operations include one or more assumptions about ploidy and / or proportions of minority species (e.g., proportions of cancer cell nucleic acids; fetal proportions). In some embodiments, plots of profiles of the processed data, further manipulated by one or more statistical and / or mathematical operations, are generated to facilitate classification and / or outcome provision. Outcomes may be provided based on plots of profiles of statistically and / or mathematically manipulated data. Outcomes presented based on plots of statistically and / or mathematically manipulated data profiles often involve one or more assumptions regarding ploidy and / or proportions of minority species (e.g., proportion of cancer cell nucleic acids; proportion of fetal species).
[0208] In some embodiments, data analysis and processing may involve the use of one or more assumptions. A suitable number or type of assumptions may be used to analyze or process a dataset. Non-limiting examples of assumptions that may be used for data processing and / or analysis include: subject ploidy, cancer cell contribution, maternal ploidy, fetal contribution, prevalence of a particular sequence in a reference population, ethnic background, prevalence of selected medical conditions in the families concerned, similarity between raw count profiles from different patients and / or between runs after GC normalization and repeat masking (e.g., GCRM), perfect match representing a PCR artifact (e.g., identical base positions), assumptions specific to nucleic acid quantification assays (e.g., fetal quantity assays (FQA)), assumptions about twins (e.g., if both twins and one twin are affected, the effective fetal proportion is only 50% of the sum of the measured fetal proportions (similarly for triplets, quadruplets, etc.)), cell-free DNA (e.g., cfDNA) uniformly covering the entire genome, and combinations thereof.
[0209] If the quality and / or depth of the mapped sequence reads does not allow for the prediction of the presence or absence of gene mutations / gene alterations and / or copy number alterations at a desired confidence level (e.g., 95% or higher) based on the normalized count profile, one or more additional mathematical manipulation algorithms and / or statistical prediction algorithms may be used to generate further numerical values useful for data analysis and / or outcome provision. The term “normalized count profile” as used herein refers to the profile generated using normalized counts. Examples of methods that may be used to generate normalized counts and normalized count profiles are described herein. As stated, mapped and counted sequence reads may be normalized with respect to the count of a test sample or a reference sample. In some embodiments, the normalized count profile may be shown as a plot.
[0210] Non-limiting examples of processing steps and normalization methods that may be used, such as normalization of windows (static or sliding), determination of weighting and bias relationships, LOESS normalization, principal component normalization, hybrid normalization, profile generation, and comparison, will be described in more detail later in this specification.
[0211] Normalization for windows (static or sliding) In certain embodiments, the processing step includes normalization to a static window, and in some embodiments, the processing step includes normalization to a moving window or a sliding window. The term “window,” as used herein, refers to one or more portions that are selected for analysis and may be used as a reference for comparison (e.g., for normalization and / or other mathematical or statistical operations). The term “normalization to a static window,” as used herein, refers to a normalization process using one or more portions selected for comparison between a subject dataset and a reference subject dataset. In some embodiments, the selected portions are used to generate a profile. A static window generally includes a predetermined set of portions that do not change during the operation and / or analysis. The terms “normalization to a moving window” and “normalization to a sliding window,” as used herein, refer to normalization performed on a portion of the genomic region of a selected test portion (e.g., an immediately adjacent portion, an adjacent portion, or a segment that encloses it), where one or more selected test portions are normalized to the portion immediately adjacent to and enclosing that selected test portion. In certain embodiments, the selected portions are used to generate a profile. Sliding window normalization or moving window normalization often involves repeatedly moving or sliding adjacent test portions and normalizing newly selected test portions to the portion immediately adjacent to or enclosing that newly selected test portion, where the adjacent window has one or more portions in common. In certain embodiments, multiple selected test portions and / or chromosomes may be analyzed by the sliding window process.
[0212] In some embodiments, normalization to a sliding or moving window may produce one or more values, where each value corresponds to normalization to a different set of reference sub-sub weighting
[0213] In some embodiments, the processing step includes weighting. The terms “weighted,” “weighting,” or “weighting function,” or their grammatical derivatives or equivalents, as used herein, refer to the mathematical operation of part or all of a dataset that may be used to alter the influence of a particular dataset feature or variable on other dataset features or variables (for example, to increase or decrease the significance and / or contribution of data contained in one or more parts or parts of the reference genome based on the quality or usefulness of the data in selected parts or parts of the reference genome). In some embodiments, a weighting function may be used to increase the influence of data with relatively small variances in the measurements and / or decrease the influence of data with relatively large variances in the measurements. For example, parts of the reference genome with under-presented or low-quality sequence data may be “lower weighted” to minimize their influence on the dataset, while selected parts of the reference genome may be “higher weighted” to increase their influence on the dataset. A non-restrictive example of a weighting function is [1 / (standard deviation) 2 ]. Partial weighting can sometimes eliminate partial dependencies. In some embodiments, one or more parts are weighted by an intrinsic function (e.g., an eigenfunction). In some embodiments, an intrinsic function involves replacing parts with orthogonal eigenparts. The weighting process may be carried out in a manner substantially similar to that of the normalization process. In some embodiments, the dataset is adjusted (e.g., by division, multiplication, addition, or subtraction) by a given variable (e.g., a weighting variable). In some embodiments, the dataset is divided by a given variable (e.g., a weighting variable). A given variable (e.g., a minimized objective function, Phi) is often chosen to weight different parts of the dataset differently (e.g., by increasing the influence of a particular data type while decreasing the influence of other data types).
[0214] Bias relationship In some embodiments, the processing step includes determining bias relationships. For example, one or more relationships are generated between local genome bias estimates and bias frequencies. The term “relationship,” as used herein, refers to a mathematical and / or graphical relationship between two or more variables or values. A relationship may be generated by a preferred mathematical and / or graphical process. Non-limiting examples of relationships include mathematical and / or graphical representations of functions, correlations, distributions, linear or nonlinear equations, lines, regressions, fitted regressions, or combinations thereof. A relationship may include a fitted relationship. In some embodiments, a fitted relationship includes a fitted regression. A relationship may include two or more weighted variables or values. In some embodiments, a relationship includes a weighted fitted regression of one or more variables or values in that relationship. A regression may be fitted in a weighted form. A regression may be fitted unweighted. In certain embodiments, generating a relationship includes plotting or graphing it.
[0215] In certain embodiments, a relationship is generated between GC density and GC density frequency. In some embodiments, a sample GC density relationship is provided by generating a relationship between (i) GC density and (ii) GC density frequency for a sample. In some embodiments, a reference GC density relationship is provided by generating a relationship between (i) GC density and (ii) GC density frequency for a reference. In some embodiments, if the local genome bias estimate is GC density, the sample bias relationship is the sample GC density relationship, and the reference bias relationship is the reference GC density relationship. The GC density in the reference GC density relationship and / or sample GC density relationship is often a presentation of local GC content (e.g., mathematical or quantitative presentation).
[0216] In some embodiments, the relationship between local genome bias estimates and bias frequencies includes a distribution. In some embodiments, the relationship between local genome bias estimates and bias frequencies includes a fitted relationship (e.g., fitted regression). In some embodiments, the relationship between local genome bias estimates and bias frequencies includes a fitted linear or nonlinear regression (e.g., polynomial regression). In certain embodiments, if local genome bias estimates and / or bias frequencies are weighted by a preferred process, the relationship between local genome bias estimates and bias frequencies includes a weighted relationship. In some embodiments, a weighted fitted relationship (e.g., weighted fit) can be obtained by a process including quantile regression, parameterized distributions, or empirical distributions with interpolation. In certain embodiments, if local genome bias estimates are weighted, the relationship between local genome bias estimates and bias frequencies for a test sample, reference, or a portion thereof includes polynomial regression. In some embodiments, a weighted fitted model includes weighting of distribution values. Distribution values may be weighted by a preferred process. In some embodiments, values located near the tails of the distribution are given smaller weights than values closer to the median of the distribution. For example, in the case of a distribution between local genome bias estimates (e.g., GC density) and bias frequencies (e.g., GC density frequencies), the weights are determined according to the bias frequencies for a given local genome bias estimate, where local genome bias estimates containing bias frequencies closer to the mean of the distribution are given larger weights than local genome bias estimates containing bias frequencies further from the mean.
[0217] In some embodiments, the processing step includes normalizing the sequence read count by comparing the local genome bias estimate of the test sample sequence reads with the local genome bias estimate of a reference (e.g., a reference genome or a portion thereof). In some embodiments, the sequence read count is normalized by comparing the bias frequency of the local genome bias estimate of the test sample with the bias frequency of the local genome bias estimate of the reference. In some embodiments, the sequence read count is normalized by comparing the sample bias relationship with the reference bias relationship, thereby generating a comparison result.
[0218] The count of sequence reads may be normalized according to the comparison results of two or more relationships. In certain embodiments, two or more relationships are compared to provide comparison results used to reduce local bias in sequence reads (e.g., normalize the count). Two or more relationships may be compared by a preferred method. In some embodiments, the comparison results include addition, subtraction, multiplication, and / or division of a first relationship with a second relationship. In certain embodiments, the comparison of two or more relationships includes the use of preferred linear and / or nonlinear regressions. In certain embodiments, the comparison of two or more relationships includes preferred polynomial regression (e.g., cubic polynomial regression). In some embodiments, the comparison results include addition, subtraction, multiplication, and / or division of a first regression with a second regression. In some embodiments, two or more relationships are compared by a process including a multiple regression inference framework. In some embodiments, two or more relationships are compared by a process including preferred multivariate analysis. In some embodiments, two or more relationships are compared by a process that includes basis functions (e.g., blending functions, e.g., polynomial basis, Fourier basis, etc.), splines, radial basis functions, and / or wavelets.
[0219] In certain embodiments, the distribution of local genome bias estimates, including bias frequencies for test samples and references, is compared by a process that includes a weighted polynomial regression of the local genome bias estimates. In some embodiments, the polynomial regression is generated between (i) a ratio (each of which includes the bias frequency of the reference local genome bias estimate and the bias frequency of the sample local genome bias estimate) and (ii) the local genome bias estimate. In some embodiments, the polynomial regression is generated between (i) the ratio of the bias frequency of the reference local genome bias estimate to the bias frequency of the sample local genome bias estimate and (ii) the local genome bias estimate. In some embodiments, the comparison of the distribution of local genome bias estimates for test samples and reference reads includes measuring the log ratio (e.g., log2 ratio) of the bias frequencies of the local genome bias estimates for reference and sample. In some embodiments, the comparison of the distribution of local genome bias estimates includes dividing the log ratio (e.g., log2 ratio) of the bias frequency of the local genome bias estimate for reference by the log ratio (e.g., log2 ratio) of the bias frequency of the local genome bias estimate for sample.
[0220] Normalizing counts according to comparison results typically involves adjusting some counts while leaving others unadjusted. Count normalization may sometimes adjust all counts, and sometimes leave no counts for sequence reads adjusted. Counts for sequence reads may be normalized by a process that includes determining weighting coefficients, which may not include directly generating and using weighting coefficients. Normalizing counts according to comparison results may sometimes involve determining weighting coefficients for each count of sequence reads. Weighting coefficients are often sequence read-specific and applied to the counts of specific sequence reads. Weighting coefficients are often determined according to the results of comparing two or more bias relationships (e.g., a reference bias relationship compared to a sample bias relationship). Normalized counts are often determined by adjusting count values according to weighting coefficients. Adjusting counts according to weighting coefficients may sometimes involve adding weighting coefficients to the counts for sequence reads, subtracting weighting coefficients from the counts for sequence reads, multiplying the counts for sequence reads by weighting coefficients, and / or dividing the counts for sequence reads by weighting coefficients. Weighting coefficients and / or normalized counts may be determined from a regression (e.g., a regression line). Normalized counts may be obtained directly from a regression line (e.g., a fitted regression line) resulting from a comparison between the bias frequencies of the local genome bias estimates of a reference (e.g., a reference genome) and the bias frequencies of the local genome bias estimates of a test sample. In some embodiments, each count of a sample read is provided with a normalized count value according to a comparison of (i) the bias frequency of the read's local genome bias estimate compared to (ii) the bias frequency of the reference local genome bias estimate. In certain embodiments, the counts of sequence reads obtained for a sample are normalized to reduce the bias in those sequence reads.
[0221] LOESS normalization In some embodiments, the processing step includes LOESS normalization. LOESS is a regression modeling method known in the art that combines multiple regression models in a metamodel based on the k-nearest neighbor method. LOESS is sometimes referred to as locally weighted polynomial regression. In some embodiments, GC LOESS applies the LOESS model to the relationship between fragment counts (e.g., sequence reads, counts) and GC composition for a portion of a reference genome. Plotting a smooth curve through a set of data points using LOESS is sometimes called a LOESS curve, particularly when each smoothed value is given by weighted quadratic least squares regression over a range of values for the criterion variable in a scatter plot on the y-axis. For each point in a given dataset, the LOESS method fits a low-order polynomial to a subset of that data, with explanatory variable values close to the point where the response is estimated. The polynomial is fitted using weighted least squares, with points closer to the point where the response is estimated being given greater weights, and points further away being given smaller weights. Next, the value of the regression function for a given point is obtained by evaluating the local polynomial using the explanatory variable values for that data point. A LOESS fit can sometimes be considered complete after the regression function value has been calculated for each data point. Many of the details of this method (e.g., the polynomial model and the degree of weighting) are flexible.
[0222] Principal component analysis In some embodiments, the processing step includes principal component analysis (PCA). In some embodiments, the sequence read count (e.g., the sequence read count of a test sample) is adjusted according to principal component analysis (PCA). In some embodiments, the lead density profile (e.g., the lead density profile of a test sample) is adjusted according to principal component analysis (PCA). The lead density profiles of one or more reference samples and / or the test subject may be adjusted according to PCA. Removing bias from the lead density profile by a PCA-related process may be referred to herein as profile adjustment. PCA may be performed by a preferred PCA method or a variation thereof. Non-limiting examples of PCA methods include canonical correlation analysis (CCA), Karhunen-Loeve transform (KLT), Hotelling transform, eigenorthogonal decomposition (POD), singular value decomposition of X (SVD), eigenvalue decomposition of XTX (EVD), factor analysis, Eckart-Young theorem, Schmidt-Mirsky theorem, empirical orthogonal functions (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic modes, spectral decomposition, empirical modal analysis, and their variations or combinations. PCA often identifies and / or adjusts one or more biases in the read density profile. Biases identified and / or adjusted by PCA are sometimes referred to herein as principal components. In some embodiments, one or more biases can be eliminated by adjusting the read density profile according to one or more principal components using a preferred method. The read density profile can be adjusted by addition, subtraction, multiplication and / or division of the read density profile with one or more principal components. In some embodiments, one or more biases may be removed from the lead density profile by subtracting one or more principal components from the lead density profile. While biases in the lead density profile are often identified and / or quantified by PCA of the profile, principal components are often subtracted from the profile at the lead density level.PCA often identifies one or more principal components. In some embodiments, PCA identifies the first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, and tenth or more principal components. In certain embodiments, one, two, three, four, five, six, seven, eight, nine, ten or more principal components are used to refine the profile. In certain embodiments, five principal components are used to refine the profile. Principal components are often used to refine the profile in the order of their appearance in PCA. For example, if three principal components are subtracted from the read density profile, the first, second, and third principal components are used. Biases identified by principal components may include profile features that are not used to refine the profile. For example, PCA may identify copy number variations (e.g., aneuploidy, minute duplication, minute deletion, deletion, translocation, insertion) and / or sex differences as principal components. Therefore, in some embodiments, one or more principal components are not used to refine the profile. For example, if the third principal component is not used to adjust the profile, the first, second, and fourth principal components may be used to adjust the profile.
[0223] The principal components can be obtained from PCA using any suitable sample or reference. In some embodiments, the principal components are obtained from a test sample (e.g., a test subject). In some embodiments, the principal components are obtained from one or more references (e.g., a reference sample, reference sequence, reference set). In certain cases, PCA is performed on the median of the read density profiles obtained from a training set containing multiple samples to identify the first and second principal components. In some embodiments, the principal components are obtained from a set of subjects lacking copy number variation of the subject. In some embodiments, the principal components are obtained from a known set of euploids. The principal components are often identified according to PCA performed using one or more read density profiles from a reference (e.g., a training set). One or more principal components obtained from the references are often subtracted from the read density profiles of the test subjects to provide an adjusted profile.
[0224] Hybrid normalization In some embodiments, the processing step includes a hybrid normalization method. The hybrid normalization method can reduce bias (e.g., GC bias) in certain cases. In some embodiments, hybrid normalization includes (i) an analysis of the relationship between two variables (e.g., count and GC content), and (ii) the selection and application of a normalization method according to that analysis. In certain embodiments, hybrid normalization includes (i) a regression (e.g., regression analysis), and (ii) the selection and application of a normalization method according to that regression. In some embodiments, the count obtained for a first sample (e.g., a first sample set) is normalized in a different way than the count obtained from another sample (e.g., a second sample set). In some embodiments, the count obtained for a first sample (e.g., a first sample set) is normalized by a first normalization method, and the count obtained from a second sample (e.g., a second sample set) is normalized by a second normalization method. For example, in a particular embodiment, the first normalization method includes the use of linear regression, and the second normalization method includes the use of nonlinear regression (e.g., LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).
[0225] In some embodiments, hybrid normalization methods are used to normalize sequence reads mapped to parts of the genome or chromosomes (e.g., counts, mapped counts, mapped reads). In certain embodiments, raw counts are normalized, and in some embodiments, adjusted, weighted, filtered, or pre-normalized counts are normalized by the hybrid normalization method. In certain embodiments, levels or Z-scores are normalized. In some embodiments, counts mapped to selected parts of the genome or chromosomes are normalized by the hybrid normalization approach. A count can refer to a preferred measure of sequence reads mapped to a part of the genome, and non-limiting examples include raw counts (e.g., unprocessed counts), normalized counts (e.g., LOESS, principal component, or normalized by a preferred method), partial levels (e.g., mean level, mean level, median level, etc.), Z-scores, etc., or combinations thereof. These counts may be raw counts or processed counts from one or more samples (e.g., test samples, pregnant woman-derived samples). In some embodiments, the count is obtained from one or more samples obtained from one or more subjects.
[0226] In some embodiments, the normalization method (e.g., type of normalization method) is selected according to regression (e.g., regression analysis) and / or correlation coefficients. Regression analysis refers to a statistical method for estimating the relationship between variables (e.g., count and GC content). In some embodiments, the regression is generated according to a measure of count and GC content for each part of multiple parts of a reference genome. A suitable measure of GC content may be used, and non-limiting examples include measures of guanine, cytosine, adenine, thymine, purine (GC) or pyrimidine (AT or ATU) content, melting temperature (T mMeasures such as (e.g., denaturation temperature, annealing temperature, hybridization temperature), free energy, or combinations thereof may be used. Measures of guanine (G), cytosine (C), adenine (A), thymine (T), purine (GC), or pyrimidine (AT or ATU) content may be expressed as a ratio or percentage. In some embodiments, any suitable ratio or percentage is used, non-limiting examples of which include GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc. or combinations thereof. In some embodiments, the measure of GC content is the ratio or percentage of GC to total nucleotide content. In some embodiments, the measure of GC content is the ratio or percentage of GC to total nucleotide content for sequence reads mapped to a portion of the reference genome. In certain embodiments, GC content is measured according to and / or from sequence reads mapped to each part of the reference genome, and these sequence reads are obtained from a sample. In some embodiments, the measure of GC content is not determined according to and / or from sequence reads. In certain embodiments, the measure of GC content is determined for one or more samples obtained from one or more subjects.
[0227] In some embodiments, regression generation includes the generation of regression analysis or correlation analysis. Suitable regressions can be used, and non-limiting examples include regression analysis (e.g., linear regression analysis), goodness-of-fit analysis, Pearson correlation analysis, rank correlation, fraction of variance unexplained, Nash-Sutcliffe model efficiency analysis, regression model validation, proportional reduction in loss, root mean square deviation, or combinations thereof. In some embodiments, a regression line is generated. In certain embodiments, regression generation includes the generation of linear regression. In certain embodiments, regression generation includes the generation of nonlinear regression (e.g., LOESS regression, LOWESS regression).
[0228] In some embodiments, the regression determines the presence or absence of a correlation (e.g., linear correlation) between, for example, a count and a measure of GC content. In some embodiments, the regression (e.g., linear regression) is generated and the correlation coefficient is determined. In some embodiments, a suitable correlation coefficient is determined, and non-limiting examples include the coefficient of determination, R 2 Examples include the Pearson correlation coefficient and other similar metrics.
[0229] In some embodiments, goodness of fit is measured for a regression (e.g., regression analysis, linear regression). Goodness of fit may be measured by visual or mathematical analysis. Evaluation may include determining whether the goodness of fit is higher for a nonlinear regression or higher for a linear regression. In some embodiments, the correlation coefficient is a measure of goodness of fit. In some embodiments, the evaluation of goodness of fit for a regression is revealed according to the correlation coefficient and / or the cutoff value of the correlation coefficient. In some embodiments, the evaluation of goodness of fit includes comparing the correlation coefficient with the cutoff value of the correlation coefficient. In some embodiments, the evaluation of goodness of fit for a regression suggests linear regression. For example, in a particular embodiment, the goodness of fit is higher for linear regression than for nonlinear regression, and the evaluation of goodness of fit suggests linear regression. In some embodiments, the evaluation suggests linear regression, and linear regression is used to normalize the counts. In some embodiments, the evaluation of goodness of fit for a regression suggests nonlinear regression. For example, in a particular embodiment, the goodness of fit is higher for nonlinear regression than for linear regression, and the evaluation of goodness of fit suggests nonlinear regression. In some embodiments, the evaluation suggests a nonlinear regression, and a nonlinear regression is used to normalize the count.
[0230] In some embodiments, the goodness-of-fit assessment suggests linear regression when the correlation coefficient is equal to or greater than the correlation coefficient cutoff. In some embodiments, the goodness-of-fit assessment suggests nonlinear regression when the correlation coefficient is less than the correlation coefficient cutoff. In some embodiments, the correlation coefficient cutoff is predetermined. In some embodiments, the correlation coefficient cutoff is about 0.5 or greater, about 0.55 or greater, about 0.6 or greater, about 0.65 or greater, about 0.7 or greater, about 0.75 or greater, about 0.8 or greater, or about 0.85 or greater.
[0231] In some embodiments, a specific type of regression is selected (e.g., linear or nonlinear regression), and after the regression is generated, the count is normalized by subtracting the regression from the count. In some embodiments, subtracting the regression from the count provides a normalized count with reduced bias (e.g., GC bias). In some embodiments, a linear regression is subtracted from the count. In some embodiments, a nonlinear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the count. Any preferred method may be used to subtract the regression line from the count. For example, if count x is derived from a portion i (e.g., portion i) containing a GC content of 0.5, and the regression line determines the count y at a GC content of 0.5, then for portion i, xy = normalized count. In some embodiments, the count is normalized before and / or after the regression subtraction. In some embodiments, the count normalized by the hybrid normalization approach is used to generate levels, Z scores, levels and / or profiles of the genome or any part thereof. In a particular embodiment, counts normalized by a hybrid normalization approach are analyzed by methods described herein to determine the presence or absence of gene mutations or gene changes (e.g., copy number changes).
[0232] In some embodiments, the hybrid normalization method includes filtering or weighting one or more portions before or after normalization. Preferred methods for filtering portions may be used, including methods for filtering portions (e.g., portions of a reference genome) as described herein. In some embodiments, portions (e.g., portions of a reference genome) are filtered before the hybrid normalization method is applied. In some embodiments, only the counts of sequencing reads mapped to selected portions (e.g., portions selected according to count variability) are normalized by hybrid normalization. In some embodiments, the counts of sequencing reads mapped to filtered reference genome portions (e.g., portions filtered according to count variability) are removed before the hybrid normalization method is used. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of a reference genome) according to preferred methods (e.g., methods described herein). In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of a reference genome) according to uncertainties for the counts mapped to each portion for multiple test samples. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to the variability of the counts. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to GC content, repeat elements, repeat sequences, introns, exons, etc., or combinations thereof. Profile
[0233] In some embodiments, the processing step includes generating one or more profiles (e.g., profile plots) from various aspects of a dataset or its differential operation (e.g., the result of one or more mathematical and / or statistical data processing steps known in the art and / or described herein).
[0234] When used herein, the term “profile” refers to the result of mathematical and / or statistical operations on data that facilitate the identification of patterns and / or correlations in large amounts of data. A “profile” often includes values resulting from one or more operations on data or datasets based on one or more criteria. A profile often includes multiple data points. Depending on the nature and / or complexity of the dataset, any suitable number of data points may be included in a profile. In certain embodiments, a profile may include two or more data points, three or more data points, five or more data points, ten or more data points, twenty-four or more data points, twenty-five or more data points, fifty or more data points, one hundred or more data points, five hundred or more data points, one thousand or more data points, five thousand or more data points, one 10,000 or more data points, or one hundred,000 or more data points.
[0235] In some embodiments, a profile represents the entire dataset, and in certain embodiments, a profile represents a part or subset of the dataset. That is, a profile may include or be generated from data points that represent data that has not been filtered to remove any data, and a profile may include or be generated from data points that represent data that has been filtered to remove unwanted data. In some embodiments, data points in a profile correspond to the results of data manipulation on a part. In certain embodiments, data points in a profile include the results of data manipulation on a group of parts. In some embodiments, groups of parts may be adjacent to each other, and in certain embodiments, groups of parts may originate from different parts of a chromosome or genome.
[0236] The data points in a profile derived from a dataset can represent any preferred categorization of the data. Non-limiting examples of categories to which data can be grouped to generate profile data points include parts based on size, parts based on sequence features (e.g., GC content, AT content, chromosomal location (e.g., short arm, long arm, centromere, telomere)), expression levels, chromosomes, or combinations thereof. In some embodiments, a profile may be generated from data points obtained from another profile (e.g., a normalized data profile renormalized to different normalization values to generate a renormalized data profile). In certain embodiments, a profile generated from data points obtained from another profile reduces the number of data points and / or the complexity of the dataset. This reduction in the number of data points and / or the complexity of the dataset often facilitates data interpretation and / or the provision of outcomes.
[0237] A profile (e.g., a genome profile, a chromosome profile, a chromosome portion profile) is often a set of normalized or unnormalized counts for two or more portions. A profile often contains at least one level, and often contains two or more levels (e.g., a profile often has multiple levels). A level is generally for a set of portions that have roughly the same count or normalized count. Levels are described in more detail herein. In certain embodiments, a profile contains one or more portions that can be weighted, removed, filtered, normalized, adjusted, averaged, derived as a mean, added, subtracted, processed, or transformed by any combination thereof. A profile often contains normalized counts mapped to portions defining two or more levels, where those counts are further normalized by a preferred method according to one of those levels. Counts in a profile (e.g., profile levels) are often related to indeterminate values.
[0238] Profiles containing one or more levels may be padded (e.g., hole padding). Padding (e.g., hole padding) refers to the process of identifying and adjusting levels in a profile that result from copy number variations (e.g., microduplications or microdeletions in the patient's genome, maternal microduplications or microdeletions). In some embodiments, levels resulting from microduplications or microdeletions in a tumor or fetus are padded. Microduplications or microdeletions in a profile may, in some embodiments, artificially increase or decrease the overall level of a profile (e.g., a chromosomal profile) that results from false-positive or false-negative determinations for chromosomal aneuploidy (e.g., trisomy). In some embodiments, levels in a profile resulting from microduplications and / or deletions are identified and adjusted (e.g., padded and / or removed) by a process sometimes referred to as padding or hole padding.
[0239] A profile containing one or more levels may contain a first level and a second level. In some embodiments, the first level is different from (e.g., significantly different from) the second level. In some embodiments, the first level contains a first subset, the second level contains a second subset, and the first subset is not a subset of the second subset. In some particular embodiments, the first subset is different from the second subset from which the first and second levels are measured. In some embodiments, a profile may have multiple first levels that are different from (e.g., significantly different from, or have significantly different values from) the second level in that profile. In some embodiments, a profile contains one or more first levels that are significantly different from the second level in that profile, and that one or more first levels are adjusted. In some embodiments, the first levels in a profile are removed from or adjusted (e.g., padded) from that profile. A profile may include multiple levels, one or more first levels that are significantly different from one or more second levels, and the majority of levels in a profile are often second levels, which are approximately equal to each other. In some embodiments, more than 50%, 60%, 70%, 80%, 90%, or 95% of the levels in a profile are second levels.
[0240] Profiles can sometimes be displayed as plots. For example, one or more levels representing partial counts (e.g., normalized counts) may be plotted and visualized. Non-limiting examples of profile plots that can be generated include raw counts (e.g., raw count profile or raw profile), normalized counts, z-scores, p-values, area ratio to fitted plicativity, median level to the ratio of fitted minority proportions to measured minority proportions, principal components, etc., or combinations thereof. In some embodiments, profile plots enable visualization of manipulated data. In certain embodiments, profile plots may be used to provide outcomes (e.g., area ratio to fitted plicativity, median level to the ratio of fitted minority proportions to measured minority proportions, principal components). The term “raw count profile plot” or “raw profile plot,” as used herein, refers to a plot of counts in each part of a region (e.g., genome, part, chromosome, chromosomal part of a reference genome, or part of a chromosome) normalized to the total counts in that region. In some embodiments, the profile may be generated using a static windowing process, and in certain embodiments, the profile may be generated using a sliding windowing process.
[0241] Profiles generated for a test subject may be compared to profiles generated for one or more reference subjects to facilitate the interpretation of mathematical and / or statistical manipulations of the dataset, and / or to provide outcomes. In some embodiments, profiles are generated based on one or more starting assumptions, e.g., assumptions described herein. In certain embodiments, test profiles often converge around predetermined values that represent the absence of copy number variation, and if a test subject has copy number variation, the values often deviate from predetermined values in regions corresponding to genomic locations where the copy number variation is located within the test subject. In test subjects at risk of or suffering from a medical condition associated with copy number variation, the values for selected regions are expected to deviate significantly from predetermined values for unaffected genomic locations. Depending on the initial assumptions (e.g., default or optimized ploidy, default or optimized cancer cell nucleic acid ratio, default or optimized fetal ratio, or a combination thereof), predetermined thresholds or cutoff values or threshold ranges that suggest the presence or absence of copy number variation may vary, but still provide an outcome useful for determining the presence or absence of copy number variation. In some embodiments, the profile suggests and / or represents the phenotype.
[0242] In some embodiments, the use of one or more reference samples substantially free of the copy number variation of the subject may be used to generate a reference count profile (e.g., a median reference count profile) which may yield a predetermined value representative of the absence of copy number variation, and if the subject has copy number variation, it will often deviate from a predetermined value in the region corresponding to the genomic location where the copy number variation is located in that subject. In subjects at risk of or suffering from a medical condition associated with copy number variation, the numerical values for the selected portion or segment are expected to deviate significantly from a predetermined value for the unaffected genomic location. In certain embodiments, the use of one or more reference samples found to have copy number variation of the subject may be used to generate a reference count profile (a median reference count profile) which may yield a predetermined value representative of the presence of copy number variation, and it will often deviate from a predetermined value in the region corresponding to the genomic location where the subject does not have copy number variation. In study subjects who are not at risk of or do not suffer from medical conditions associated with copy number variation, the values for the selected portion or segment are expected to differ significantly from the predetermined values for the affected genomic location.
[0243] As a non-limiting example, normalized sample count profiles and / or normalized reference count profiles can be obtained from raw sequence read data by (a) calculating the median reference count for a selected chromosome, part or portion thereof from a set of references known to have no copy number variation; (b) removing (e.g., filtering) portions of non-informational value from the raw counts of the reference sample; (c) normalizing the reference counts for all remaining portions of the reference genome against the total number of remaining counts for the selected chromosome or selected genomic location of the reference sample (e.g., the sum of the counts remaining after removing portions of non-informational value from the reference genome), thereby generating a normalized reference subject profile; (d) removing the corresponding portions from the test subject sample; and (e) normalizing the remaining test subject counts for one or more selected genomic locations against the sum of the median remaining reference counts for the chromosome containing the selected genomic location, thereby generating a normalized test subject profile. In certain embodiments, a further normalization step for the entire genome, which is reduced by the filtered portions in (b), may be included between (c) and (d).
[0244] In some embodiments, a lead density profile is measured. In some embodiments, the lead density profile includes at least one lead density, and often includes two or more lead densities (e.g., a lead density profile often includes multiple lead densities). In some embodiments, the lead density profile includes a suitable quantitative value (e.g., mean, median, Z-score, etc.). The lead density profile often includes values resulting from one or more lead densities. The lead density profile may include values resulting from one or more operations on the lead density based on one or more adjustments (e.g., normalization). In some embodiments, the lead density profile includes an unoperated lead density. In some embodiments, one or more lead density profiles are generated from various aspects of a dataset including lead densities or their derivatives (e.g., the result of one or more mathematical and / or statistical data processing steps known in the art and / or described herein). In certain embodiments, the lead density profile includes a normalized lead density. In some embodiments, the lead density profile includes an adjusted lead density. In certain embodiments, a lead density profile may include raw lead density (e.g., unmanipulated, unadjusted, or unnormalized lead density), normalized lead density, weighted lead density, filtered portion of lead density, z-score of lead density, p-value of lead density, integral of lead density (e.g., area under the curve), mean, mean or median, principal components, or a combination thereof. The lead density and / or lead density profile of a lead density profile are often related to a measure of uncertainty (e.g., MAD). In certain embodiments, a lead density profile may include a distribution of median lead densities. In some embodiments, a lead density profile may include relationships between multiple lead densities (e.g., fitting relationships, regressions, etc.).For example, a read density profile may include a relationship between read density (e.g., read density values) and genomic location (e.g., location of a part). In some embodiments, the read density profile is generated using a static windowing process, and in certain embodiments, the read density profile is generated using a sliding windowing process. In some embodiments, the read density profile may be printed and / or displayed (e.g., visually displayed, e.g., as a plot or graph).
[0245] In some embodiments, the read density profile corresponds to a partial set (e.g., a partial set of a reference genome, a partial set of a chromosome, or a partial subset of a part of a chromosome). In some embodiments, the read density profile includes read density and / or read count associated with a set of parts (e.g., a set, a subset). In some embodiments, the read density profile is measured against the read density of a contiguous portion. In some embodiments, a contiguous portion includes a region of the reference sequence and / or gaps containing sequence reads not included in the density profile (e.g., portions removed by filtering). A contiguous portion (e.g., a partial set) may correspond to adjacent regions of the genome or adjacent regions of a chromosome or gene. For example, two or more contiguous portions may be a sequence assembly of DNA sequences longer than each portion when aligned by merging those portions end-to-end. For example, two or more contiguous portions may be an intact genome, chromosome, gene, intron, exon, or a portion thereof. A read density profile may be determined from a set of contiguous and / or non-contiguous portions (e.g., a set, a subset). In some cases, a read density profile may contain one or more parts that can be weighted, removed, filtered, normalized, adjusted, averaged, derived as an average, added, subtracted, processed, or transformed by any combination thereof.
[0246] Read density profiles are often measured against a sample and / or reference (e.g., a reference sample). Read density profiles may be generated for the entire genome, one or more chromosomes, or a portion of the genome or chromosomes. In some embodiments, one or more read density profiles are measured against the genome or a portion thereof. In some embodiments, a read density profile represents the entire set of read densities in a sample, and in certain embodiments, a read density profile represents a portion or subset of the read densities in a sample. That is, a read density profile may include or be generated from read densities representing data that has not been filtered to remove any unwanted data, and a read density profile may include or be generated from data points representing data that has been filtered to remove unwanted data.
[0247] In some embodiments, the read density profile is measured against a reference (e.g., a reference sample, a training set). The read density profile against a reference may be referred to herein as the reference profile. In some embodiments, the reference profile includes the read density obtained from one or more references (e.g., a reference sequence, a reference sample). In some embodiments, the reference profile includes the read density measured against one or more known euploidy samples (e.g., a set of known euploidy samples). In some embodiments, the reference profile includes the read density of a filtered portion. In some embodiments, the reference profile includes the read density adjusted according to one or more principal components.
[0248] Conducting a comparison In some embodiments, the processing step includes a preforming step (e.g., a step of comparing a test profile with a reference profile). Two or more datasets, two or more relationships, and / or two or more profiles may be compared by a preferred method. Non-exclusive examples of statistical methods suitable for comparing datasets, relationships, and / or profiles include the Behrens-Fisher approach, bootstrapping, Fisher's method for combining independent significance tests, Neyman-Pearson test, confirmatory data analysis, exploratory data analysis, exact tests, F-tests, Z-tests, T-tests, calculation and / or comparison of uncertainty measures, null hypotheses, alternative null hypotheses, etc., chi-squared tests, summative tests, calculation and / or comparison of significance levels (e.g., statistical significance levels), meta-analysis, multivariate analysis, regression, linear simple regression, robust linear regression, etc., or a combination of the above. In certain embodiments, the comparison of two or more datasets, relationships, and / or profiles includes the measurement and / or comparison of uncertainty measures. When used herein, “measure of uncertainty” refers to a measure of significance (e.g., statistical significance), a measure of error, a measure of variance, a measure of confidence, or a combination thereof. A measure of uncertainty can be a value (e.g., a threshold) or a range of values (e.g., an interval, a confidence interval, a Bayesian confidence interval, a threshold range). Non-restrictive examples of measures of uncertainty include p-values, measures of preferred deviation (e.g., standard deviation, sigma, absolute deviation, mean absolute deviation, etc.), measures of preferred error (e.g., standard error, mean square error, root mean square error, etc.), measures of preferred variance, preferred standard scores (e.g., standard deviation, cumulative percentage, percentile equivalent, Z-score, T-score, R-score, standard nine (StanNine), percentage in StanNine, etc.), or a combination thereof. In some embodiments, determining the significance level includes determining a measure of uncertainty (e.g., p-value).In certain embodiments, two or more datasets, relationships, and / or profiles may be analyzed and / or compared by using multiple (e.g., two or more) statistical methods (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, K-nearest neighbors, logistic regression, and / or loss smoothing) and / or any suitable mathematical and / or statistical operations (e.g., referred to herein as operations).
[0249] In some embodiments, the processing step includes comparing two or more profiles (e.g., two or more read density profiles). The comparison of profiles may include comparing profiles generated for selected regions of the genome. For example, if the test profile and the reference profile are measured for a region of the genome (e.g., a reference genome) that is substantially the same region, the test profile may be compared to the reference profile. The comparison of profiles may sometimes include comparing two or more subsets of parts of a profile (e.g., a read density profile). A subset of parts of a profile may correspond to a region of the genome (e.g., a chromosome or a region thereof). A profile (e.g., a read density profile) may contain any number of subsets of parts. A profile (e.g., a read density profile) may contain two or more, three or more, four or more, or five or more subsets. In a particular embodiment, if each part is a region of an adjacent reference genome, the profile (e.g., a read density profile) contains two subsets of parts. In some embodiments, if both the test profile and the reference profile include a first subset of the part and a second subset of the part, and the first and second subsets are different regions of the genome, then the test profile can be compared to the reference profile. Some subsets of the part of the profile may contain copy number variations, while other subsets of the part may substantially not contain copy number variations. It may be the case that all subsets of the part of the profile (e.g., the test profile) substantially do not contain copy number variations. It may be the case that all subsets of the part of the profile (e.g., the test profile) contain copy number variations. In some embodiments, the test profile may include a first subset of the part containing copy number variations and a second subset of the part substantially not containing copy number variations.
[0250] In certain embodiments, the comparison of two or more profiles involves determining and / or comparing a measure of uncertainty for two or more profiles. Profiles (e.g., read density profiles) and / or associated measures of uncertainty may be compared to facilitate the interpretation of mathematical and / or statistical operations on the dataset, and / or to provid...
Claims
1. A method implemented by a computer, A step of obtaining a test file for a test subject using a computing system, wherein the test file includes alignment information of sequence reads generated by sequencing a test sample obtained from the test subject, and the test sample includes multiple nucleic acid species and a small number of nucleic acid species; A step of determining a minority ratio in a test sample using the test file, wherein the minority ratio is the ratio of the quantitative value of a minority nucleic acid to the quantitative value of the total nucleic acid; A step of using the test file to filter the alignment information using the computing system to obtain filtered alignment information; A step of segmenting the target region of the genome of the test subject using the computing system based on the filtered alignment information, wherein the step of performing the segmentation is: To determine the type of segmentation to be implemented based on the aforementioned minority ratio. The target region is segmented using concentrated window segmentation if the aforementioned minority ratio is less than the threshold, or using genome-wide segmentation if the aforementioned minority ratio is above or equal to the threshold. A process of performing segmentation, including; and A step of determining, based on the segmentation, the presence or absence of a gene mutation in the test subject by the computing system, wherein the gene mutation is a duplication, deletion, fusion, insertion, short tandem repeat (STR), mutation, single nucleotide change, rearrangement, substitution or abnormal methylation. Methods that include...
2. The computer-implemented method according to claim 1, wherein the threshold is 10% to 12%.
3. The step further includes identifying a predetermined set of genomic subsets in the target region, wherein the identifying step is To obtain multiple candidate subregions within the target region; Obtaining one or more precision measures for each of the multiple candidate subregions for multiple samples in a training set, wherein each of the multiple samples is classified as having copy number variation in the target region; Selecting a subregion for the predetermined set of genomic subsets that provides optimal accuracy based on one or more of the aforementioned accuracy measures. A computer-implemented method according to claim 1, including the method described in claim 1.
4. The computer-implemented method according to claim 3, wherein each genome segment is 1 megabase to 40 megabases in length.
5. The computer-implemented method according to claim 1, wherein the minority ratio is a cancer ratio, a tumor ratio, or a fetal ratio.
6. The computer-implemented method according to claim 1, wherein, if the minority ratio exceeds or is equal to the threshold, the genome-wide segmentation is circular binary segmentation (CBS), wavelet segmentation, Fourier transform-based segmentation, sliding window segmentation, Markov strand model-based segmentation, maximum entropy segmentation, binary recursive segmentation, or level-based segmentation.
7. The computer-implemented method according to claim 1, wherein the segmentation is performed by repeatedly dividing the target region into regions of equal copy number using a likelihood ratio statistic.
8. The computer-implemented method according to claim 1, wherein the filtering is performed based on a quality score associated with each sequence read.
9. A step of the computing system determining a first set of scores corresponding to the total number of sequence reads mapped to a set of bins based on the filtered alignment information; and The computing system performs the process of normalizing the first set of scores to generate a normalized first set of scores. The computer-implemented method according to claim 1, further comprising:
10. The method implemented on a computer according to claim 9, wherein the normalization is self-normalization.
11. The method implemented on a computer according to claim 9, further comprising the step of obtaining a reference file containing alignment information for a reference sample or a set of reference samples using the computing system, wherein the normalization is based on a comparison between the test file and the reference file.
12. The computer-implemented method according to claim 1, wherein the presence or absence of the gene mutation is determined based on a predetermined threshold.
13. The computer-implemented method according to claim 1, further comprising the step of sequencing the test sample to obtain the sequence reads, wherein the sequencing is whole-genome sequencing or targeted sequencing.
14. The computer-implemented method according to claim 1, wherein the test file further includes array features based on the alignment information, and the segmentation is further determined based on the array features.
15. The computer-implemented method according to claim 13, wherein the depth of the array is 0.01 to 100 times.
16. The computer-implemented method according to claim 1, wherein the filtering includes filtering out sequence reads mapped to repeat regions or low-mapping regions.
17. The computer-implemented method according to claim 1, further comprising the step of performing GC bias correction based on the alignment information or the filtered alignment information before determining the segmentation.
18. The process of data mining a disease database to determine a set of genomic regions; A step of aligning the sequence reads to the set of genomic regions of a reference genome to generate the alignment information; and A step of retaining sequence reads aligned to a subset of the set of genomic regions by filtering out one or more portions of the set of genomic regions based on predetermined criteria. The computer-implemented method according to claim 1, further comprising:
19. The computer-implemented method according to claim 1, wherein the step of determining the presence or absence of the gene mutation includes determining the presence or absence of copy number changes in the minority nucleic acid species based on the sequence reads, the filtered alignment information, and the segmentation.
20. The computer-implemented method according to claim 19, wherein the copy number change is aneuploidy, duplication of one or more chromosomes, loss of one or more chromosomes, partial chromosomal abnormality, mosaicism, translocation, or inversion.
21. The computer-implemented method according to claim 1, wherein the minority ratio is determined based on the ratio of alleles of the polymorphic sequence, and the polymorphic sequence is determined based on the filtered alignment information.
22. With one or more data processors; A non-temporary computer-readable storage medium, which, when executed on one or more data processors, includes an instruction to cause one or more data processors to perform the method according to any one of claims 1 to 21, and A system that includes this.
23. A non-temporary computer-readable storage medium storing programmed instructions, wherein when the programmed instructions are executed by a computer processor, the computer processor causes the computer processor to perform the method according to any one of claims 1 to 21.