Detection of mutations and ploidy of chromosome segments
The method and system improve the detection of chromosome ploidy and SNVs by correcting errors and estimating phase to accurately identify CNVs and SNVs, facilitating early disease diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2026-03-27
AI Technical Summary
Current methods are inadequate for accurately detecting copy number variations (CNVs) and single nucleotide variants (SNVs) in chromosomes, which are associated with various diseases and conditions, including cancer, mental disorders, and autoimmune diseases, necessitating improved diagnostic methods for early detection and treatment.
A method and system for determining chromosome ploidy by receiving allele frequency data, correcting errors, estimating phase, and creating joint probabilities to select the optimal model for chromosomal ploidy, and detecting single nucleotide variants by analyzing nucleic acid sequences and correcting amplification biases.
Enhances the accuracy of detecting chromosome segment deletions and duplications, enabling early diagnosis of diseases, identifying disease risk, and detecting circulating tumor nucleic acids, with a detection limit as low as 0.015% for SNVs.
Smart Images

Figure 0007836659000030 
Figure 0007836659000031 
Figure 0007836659000032
Abstract
Description
[Technical Field]
[0001] References to related applications This application claims priority to U.S. Provisional Patent Application No. 61 / 982,245 filed on 21 April 2014, U.S. Provisional Patent Application No. 61 / 987,407 filed on 1 May 2014, U.S. Provisional Patent Application No. 62 / 066,514 filed on 21 October 2014, U.S. Provisional Patent Application No. 62 / 146,188 filed on 10 April 2015, U.S. Provisional Patent Application No. 62 / 147,377 filed on 14 April 2015, and U.S. Provisional Patent Application No. 62 / 148,173 filed on 15 April 2015, the teachings of these applications being incorporated herein by reference in their entirety.
[0002] The present invention relates primarily to a method and system for detecting the ploidy of chromosome segments, and a method and system for detecting single nucleotide variants. [Background technology]
[0003] Copy number diversity (CNV) is recognized as the primary cause of structural diversity in the genome, typically involving sequence duplications and deletions ranging in length from 1,000 base pairs (1 kb) to 20 megabase pairs (mb). Deletions and duplications of chromosome segments or entire chromosomes are associated with a variety of conditions, including susceptibility to or resistance to disease.
[0004] CNVs are often classified into one of two main types based on the length of the affected sequences. The first type includes copy number variations (CNPs). CNPs are common in the general population and occur at an overall frequency of over 1%. CNPs are usually small (mostly less than 10 kilobase pairs in length) and are common in genes encoding proteins important for drug detoxification and immunity. Subsets of these CNPs have vastly different copy numbers. As a result, the chromosomal copy number for a particular set of genes can vary greatly from person to person (e.g., 2, 3, 4, 5, etc.). CNPs associated with immune response genes have recently been linked to susceptibility to rare hereditary diseases (including psoriasis, Crohn's disease, and glomerulonephritis).
[0005] The other type of CNV includes relatively rare variants that are considerably longer than CNPs, ranging in length from approximately 100,000 base pairs to over 1 million base pairs. In some cases, these CNVs may have arisen during the formation of sperm or eggs that gave rise to a particular individual, or may have been inherited only within a few generations of a family. Such rare large structural variants have been observed more frequently in subjects with intellectual disability, developmental delay, schizophrenia, and autism. Based on the characteristics of these subjects, it is hypothesized that these rare large CNVs may be more important than other forms of inherited mutations (including single nucleotide substitutions) in neurocognitive disorders.
[0006] Gene copy numbers may be altered in cancer cells. For example, Chr1p duplication is commonly seen in breast cancer. Also, EGFR copy numbers can be higher in non-small cell lung cancer compared to normal cells. Cancer is one of the leading causes of death, and early diagnosis and treatment of cancer are important because they can improve patient outcomes (for example, by improving the probability and duration of sedation). Early diagnosis can also reduce the amount or number of treatments a patient receives. Many current treatments that destroy cancerous cells also affect normal cells, which can result in various side effects (such as nausea, vomiting, decreased blood cell counts, increased risk of infection, hair loss, and mucosal ulcers). Therefore, early detection of cancer is desirable because it can reduce the number and / or amount of treatments (such as chemotherapy drugs or radiation).
[0007] Copy number variations have also been linked to several mental, physical, and idiopathic learning disorders. Non-invasive prenatal testing (NIPT) using cell-free DNA can detect abnormalities (e.g., trichromacy, triploidy, and sex aneuploidy of fetal chromosomes 13, 18, and 21). Subchromosomal microdeletions can also cause several mental and physical disorders, but are more difficult to detect due to their smaller size. Eight of the microdeletion syndromes have agglutination rates exceeding 1 in 1000, similar to the incidence of autosomal trichromacy in the fetus.
[0008] Furthermore, a high copy number of CCL3L1 has been associated with lower susceptibility to HIV infection. A low copy number of FCGR3B (CD16 cell surface immunoglobulin receptor) may increase susceptibility to systemic lupus erythematosus and similar inflammatory autoimmune diseases.
[0009] Thus, there is a need for improved methods to detect deletions and duplications of chromosome segments or entire chromosomes. It is desirable that these methods can be used to accurately diagnose diseases such as cancer, identify increased risk of disease development, or diagnose fetal CNVs. [Overview of the project]
[0010] In exemplary embodiments of this specification, a method is provided for determining the ploidy of chromosome segments contained in a sample of an individual. The method includes the following steps: a. Allele frequency data is received, which includes the amount of each allele present at each locus in the set of polymorphic loci within the chromosome segment in the sample. b. By estimating the phase of the allele frequency data, phased allele information for the set of polymorphic loci is created. c. Using the allele frequency data, create individual probabilities of the allele frequencies of the polymorphic locus for different ploidy states. d. Using the individual probabilities and the phase-determined allele information, create a joint probability for the set of polymorphic loci. e. Determining the ploidy of a chromosome segment by selecting the optimal model that exhibits chromosomal ploidy based on the joint probability.
[0011] In one exemplary embodiment of the method for determining ploidy, the data is prepared using nucleic acid sequence data, particularly high-processing nucleic acid sequence data. In one example of the method for determining ploidy, errors are corrected before the allele frequency data is used to prepare individual probabilities. In certain exemplary embodiments, the errors corrected include bias in allele amplification efficiency. In other embodiments, the errors corrected include environmental impurities and impurities from genotyping analysis. In some embodiments, the errors corrected include bias in allele amplification, environmental impurities, and impurities from genotyping analysis.
[0012] In certain embodiments of the method for determining ploidy, the individual probabilities are constructed using a set of models for both different ploidy states and allele imbalance rates relating to the set of polymorphic loci. In these embodiments and other embodiments, the joint probabilities are constructed by considering linkages between polymorphic loci in the chromosome segment.
[0013] Accordingly, in an exemplary embodiment combining some of these embodiments, a method for detecting chromosome ploidy in an individual sample is provided herein. The method comprises the following steps: a. Receiving nucleic acid sequence data relating to alleles present in a set of polymorphic loci in the chromosomal segments of the individual, b. Using the nucleic acid sequence data, detect the allele frequency in the set of polymorphic loci, c. Correct the bias in the allele amplification efficiency of the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci, d. By estimating the phase of the nucleic acid sequence data, phase-determined allele information relating to the set of polymorphic loci is created. e. By comparing the modified allele frequencies with a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci, we can create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states. f. By combining the individual probabilities while considering the linkage between polymorphic loci in the chromosome segment, a joint probability for the set of polymorphic loci is created. g. Select the optimal model that exhibits chromosomal aneuploidy based on the joint probability.
[0014] In other aspects of this specification, a system for detecting chromosome ploidy of an individual sample, the system is a. An input processor configured to receive allele frequency data including the amount of each allele present at each locus in the set of polymorphic loci in the chromosome segment in the sample, b. Modeler, i. By estimating the phase of the allele frequency data, phase-determined allele information relating to the set of polymorphic loci is created. ii. Using the allele frequency data, create individual probabilities of the allele frequency of the polymorphic locus for different ploidy states. iii. A modeler configured to create a joint probability for the set of polymorphic loci using the individual probabilities and the phase-determined allele information, c. A hypothesis manager configured to determine the ploidy of a chromosome segment by selecting the optimal model that exhibits chromosomal ploidy based on the joint probability.
[0015] In a particular embodiment of this system, the allele frequency data is data generated by a nucleic acid sequencing system. In a particular embodiment, the system further includes an error correction unit configured to correct errors in the allele frequency data, and the corrected allele frequency data is used by the modeler to create individual probabilities. In a particular embodiment, the error correction unit corrects bias in allele amplification efficiency. In a particular embodiment, the modeler creates the individual probabilities using a set of models for both different ploidy states and allele imbalance rates relating to the set of polymorphic loci. In a particular exemplary embodiment, the modeler creates the joint probabilities by considering linkage between polymorphic loci in the chromosome segment.
[0016] In this specification, in one exemplary embodiment, a system for detecting chromosome ploidy in an individual sample is provided. The system is a. An input processor configured to receive nucleic acid sequence data relating to alleles present in a set of polymorphic loci in a chromosomal segment of the individual, and to detect the allele frequency in the set of polymorphic loci using the nucleic acid sequence data, b. An error correction unit configured to correct errors in the detected allele frequencies and create corrected allele frequencies for the set of polymorphic loci, c. Modeler, i. By estimating the phase of the nucleic acid sequence data, phase-determined allele information relating to the set of polymorphic loci is created. ii. By comparing the phase-determined allele information with a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci, we can create individual probabilities of the allele frequencies of the polymorphic loci for different ploidy states. iii. A modeler configured to create a joint probability for a set of polymorphic loci by combining the individual probabilities while considering the relative distances between the polymorphic loci in the chromosome segment, d. Includes a hypothesis management unit configured to select the optimal model showing chromosomal aneuploidy based on the joint probability.
[0017] In one aspect, the present invention provides a method for determining whether circulating tumor nucleic acids are present in a sample of an individual. The method is a. Analyze the sample to determine the ploidy of the set of polymorphic loci in the chromosome segments of the individual, b. The process includes determining the level of allele imbalance present at the polymorphic locus based on the ploidy determination, where an allele imbalance rate of 0.4%, 0.45%, or 0.5% or higher indicates the presence of circulating tumor nucleic acids in the sample.
[0018] In certain embodiments, a method for determining the presence of circulating tumor nucleic acids further comprises detecting single nucleotide variants of single nucleotide variant positions included in a set of single nucleotide variant positions, wherein the detection of an allele mismatch rate of 45% or higher and / or the detection of single nucleotide variants indicates the presence of circulating tumor nucleic acids in the sample.
[0019] In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing 1,000 to 50,000 or 100 to 1,000 polymorphic loci for ploidy.
[0020] In this specification, a method for detecting single nucleotide variants in a sample is provided in certain respects. Accordingly, a method for determining whether single nucleotide variants are present in a set of genomic locations contained in a sample of an individual is provided. The method is a. For each genome location, use the training dataset to create estimates of the efficiency and error rate per cycle for amplification products across that genome location. b. Obtaining observed nucleotide identity information regarding the location of each genome contained in the sample, c. Using the estimated amplification efficiency and error rate per cycle for each genome location individually, the set of probabilities of single nucleotide variants resulting from one or more actual mutations at each genome location is determined by comparing the observed nucleotide identity information for each genome location with models of different variant rates. d. This includes determining the most likely real variety rate and confidence level from the set of probabilities for each genome location.
[0021] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the estimates of efficiency and error rate per cycle are made with respect to a series of amplification products across the genomic location. For example, this could include 2, 3, 4, 5, 10, 15, 20, 25, 50, 100 or more amplification products across the genomic location. In a particular embodiment, the detection limit in the method for detecting one or more SNVs is 0.015%, 0.017%, or 0.02%.
[0022] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the observed nucleotide identity information includes the observed total number of reads for each genomic location and the observed number of reads for each diversity allele for each genomic location.
[0023] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the sample is a plasma sample and the single nucleotide variant is present in the circulating tumor DNA of the sample.
[0024] In other embodiments of this specification, a method is provided for detecting one or more single nucleotide variants present in a test sample of an individual. The method according to this embodiment includes the following steps: a. Based on the results generated during sequencing, the median diversity allele frequency of multiple control samples obtained from each of multiple normal individuals is determined for each single nucleotide diversity position included in the set of single nucleotide diversity positions. This identifies selected single nucleotide diversity positions that have a median diversity allele frequency of normal samples that is below the threshold. After excluding outlier samples for each of these single nucleotide diversity positions, the background error for each of these single nucleotide diversity positions is determined. b. Based on the data generated during sequencing of the test sample, the weighted mean and variance based on the observed read depth are determined for the selected single nucleotide diversity position of the test sample. c. Detecting one or more single nucleotide variants by using a computer to identify one or more single nucleotide variant locations in which the weighted average based on read depth is statistically significant compared to the background error.
[0025] In a particular embodiment, in the method for detecting one or more SNVs, the sample is a plasma sample, the control sample is a plasma sample, and the detected one or more single nucleotide variants are present in the circulating tumor DNA of the sample. In a particular embodiment, in the method for detecting one or more SNVs, the plurality of control samples include at least 25 samples. In a particular embodiment, in the method for detecting one or more SNVs, the weighted average and observed variance by observed read depth are calculated by excluding outliers from the data created when performing high-process sequencing. In a particular embodiment, in the method for detecting one or more SNVs, the read depth for each single nucleotide variant position in the test sample is at least 100 reads.
[0026] In certain embodiments, the method for detecting one or more SNVs includes a multiple amplification reaction performed under restrictive primer reaction conditions, wherein the sequencing is performed under restrictive primer reaction conditions. In certain embodiments, the method for detecting one or more SNVs has a detection limit of 0.015%, 0.017%, or 0.02%.
[0027] In one aspect, the present invention is characterized by a method for determining whether there is an over-recurrence of copy numbers in a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells derived from a given individual. In some embodiments, the method includes obtaining phase-determined genetic data of the first homologous chromosome segment, including the identity of the allele present at each locus in the first homologous chromosome segment, for each locus in the set of polymorphic loci of the first homologous chromosome segment; obtaining phase-determined genetic data of the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the second homologous chromosome segment; and obtaining measured allele genetic data, including the amount of each allele present in a sample of DNA or RNA derived from one or more cells derived from the given individual, for each allele present at each locus in the set of polymorphic loci. In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment in the genome of one or more cells derived from the individual; calculating the likelihood of the one or more hypotheses (by computer, etc.) based on the obtained sample genetic data and phase-determined genetic data; selecting the most likely hypothesis to determine the degree of overpopulation of the copy number of the first homologous chromosome segment in the genome of one or more cells derived from the individual. In some embodiments, the phase-determined data includes population-based haplotype frequencies and / or phase-determined data inferred using measured phase-determined data (e.g., phase-determined data obtained by measuring a sample containing DNA or RNA derived from the individual or a relative of the individual).
[0028] In one aspect, the present invention provides a method for determining whether there is an over-recurrence of copy numbers in a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells derived from a given individual. In some embodiments, the method includes obtaining phase-determined genetic data of the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the first homologous chromosome segment; obtaining phase-determined genetic data of the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the second homologous chromosome segment; and obtaining measured allele genetic data, including the amount of each allele present in a sample of DNA or RNA derived from one or more cells derived from the given individual, for each allele present at each locus in the set of polymorphic loci. In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment; calculating expected gene data for multiple loci of the sample from phase-determined gene data obtained for each of the hypotheses; calculating the data fit between the gene data of the sample obtained for the sample and the expected gene data (using a computer, etc.); ranking one or more of the hypotheses according to the data fit; selecting the highest-ranked hypothesis to determine the degree of overpopulation of the copy number of the first homologous chromosome segment in the genome of one or more cells derived from the individual.
[0029] In one aspect, the present invention is characterized by a method for determining whether there is an over-reproduction of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells of an individual. In some embodiments, the method includes obtaining phase-determined genetic data of the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the first homologous chromosome segment; obtaining phase-determined genetic data of the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the second homologous chromosome segment; and obtaining measured allele genetic data, including the amount of each allele present in a DNA or RNA sample derived from one or more target cells and one or more non-target cells from the individual, for each allele present at each locus in the set of polymorphic loci. In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment; for each of the hypotheses, calculating (by computer, etc.) expected gene data for the multiple loci in the sample for one or more possible ratios of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample from the obtained phase-determined gene data; calculating (by computer, etc.) the data fit between the obtained gene data of the sample and the expected gene data of the sample for each possible ratio of DNA or RNA and hypothesis; ranking one or more of the hypotheses according to the data fit; selecting the highest-ranked hypothesis to determine the degree of overpopulation of the copy number of the first homologous chromosome segment of the genome of one or more cells derived from the individual.
[0030] In one aspect, the present invention is characterized by a method for determining whether there is an over-reproduction of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells derived from a given individual. In some embodiments, the method includes obtaining phase-determined genetic data of the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the first homologous chromosome segment; obtaining phase-determined genetic data of the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the second homologous chromosome segment; and obtaining measured allele genetic data, including the amount of each allele present in a DNA or RNA sample derived from one or more target cells and one or more non-target cells derived from the given individual, for each allele present at each locus in the set of polymorphic loci. In some embodiments, the method includes: listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment; for each of the hypotheses, calculating expected gene data for the multiple loci in the sample for one or more possible ratios of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample from the obtained phase-determined gene data (using a computer, etc.); for each of the multiple loci, the possible ratio of DNA or RNA, and the hypothesis, calculating the likelihood that the hypothesis is correct by comparing the obtained gene data from the sample for that locus with the possible ratio of DNA or RNA and the expected gene data for that locus under the hypothesis (using a computer, etc.); determining the composite probability of each hypothesis by combining the probabilities of each locus in the hypothesis for each locus and possible ratio; selecting the hypothesis with the highest composite probability to determine the degree of overpopulation of the copy number of the first homologous chromosome segment. In some embodiments, all polymorphic loci are considered simultaneously to calculate the probability of a particular hypothesis, and the hypothesis with the highest probability is selected.
[0031] In one aspect, the present invention features a method for determining the copy number of a target chromosome segment in the genome of a fetus. In some embodiments, the method comprises obtaining phase-determined genetic data of at least one biological parent of the fetus, wherein the phase-determined genetic data comprises the identity of alleles present at each locus of a set of polymorphic loci in a homologous chromosome segment pair containing the target chromosome segment, specifically in the first homologous chromosome segment and the second homologous chromosome segment. In some embodiments, the method comprises obtaining genetic data of a set of polymorphic loci in the target chromosome segment by measuring the amount of each allele present at each locus in a DNA or RNA mixture sample containing fetal DNA or RNA and maternal DNA or RNA derived from the mother of the fetus. In some embodiments, the method comprises listing one or more sets of hypotheses that define the copy number of the target chromosome segment present in the genome of the fetus. In some embodiments, the method includes listing a set of one or more hypotheses that define, for one or both parents, the copy number of a first homologous chromosome segment or portion derived from that parent in the fetal genome, the copy number of a second homologous chromosome segment or portion derived from that parent in the fetal genome, and the total copy number of the target chromosome segment present in the fetal genome. In some embodiments, the method includes, for each of the hypotheses, calculating (by computer, etc.) expected gene data for multiple loci in the mixed sample from the obtained phase-determined gene data from the (both) parents, calculating (by computer, etc.) the data fit between the obtained gene data of the mixed sample and the expected gene data of the mixed sample, ranking one or more of the hypotheses based on the data fit, selecting the highest-ranked hypothesis, and determining the copy number of the target chromosome segment in the fetal genome.
[0032] In one aspect, the present invention features a method for determining the copy number of a target chromosome or chromosome segment in the genome of a fetus. In some embodiments, the method comprises obtaining phase-determined genetic data for at least one of the biological parents of the fetus, wherein the phase-determined genetic data comprises the identity of alleles present at each locus of a set of polymorphic loci in a first homologous chromosome segment and a second homologous chromosome segment of the parent. In some embodiments, the method comprises obtaining genetic data of a set of polymorphic loci in a chromosome or chromosome segment by measuring the amount of each allele present at each locus in a DNA or RNA mixture sample containing fetal DNA or RNA and maternal DNA or RNA from the fetus's mother. In some embodiments, the method comprises listing one or more sets of hypotheses that define the copy number of a target chromosome or chromosome segment present in the genome of the fetus. In some embodiments, the method includes, for each of the hypotheses, creating a probability distribution of the expected amount of each allele present at each of the multiple loci in the mixed sample (e.g., created on a computer) from (i) the obtained phase-determined gene data from the (both) parents and optionally (ii) the probability of one or more crossovers that may have occurred during gamete formation that resulted in a copy of the target chromosome or chromosome segment in the fetus; for each of the hypotheses, calculating (by computer, etc.) the fit between (1) the obtained gene data of the mixed sample and (2) the probability distribution of the expected amount of each allele at each of the multiple loci in the mixed sample for that hypothesis; ranking one or more of the hypotheses based on the data fit; selecting the highest-ranked hypothesis; and determining the copy number of the target chromosome segment in the fetal genome.
[0033] In some embodiments, the method includes obtaining phase-determined genetic data of the mother of the fetus. In some embodiments, the method includes listing a set of one or more hypotheses that define the copy number of a first maternal homologous chromosome segment or portion thereof in the fetal genome, the copy number of a second maternal homologous chromosome segment or portion thereof in the fetal genome, and the total copy number of the target chromosome segment present in the fetal genome. In some embodiments, the method includes, for each of the hypotheses, calculating expected genetic data for a plurality of loci in the mixed sample from the obtained maternal phase-determined genetic data.
[0034] In some embodiments, the expected genetic data for each of the hypotheses includes the identity and quantity of one or more alleles present at each of multiple loci in the maternal DNA or RNA and fetal DNA or RNA in the mixed sample. In some embodiments, the method includes calculating the expected genetic data (by computer, etc.) by determining the fetal DNA or RNA fraction and the maternal DNA or RNA fraction in the mixed sample. In some embodiments, the method includes, for each of the multiple loci, calculating the expected quantity of one or more alleles for that locus in the maternal DNA or RNA in the mixed sample, using the identity of the alleles present at that locus in the obtained maternal phase-determined genetic data and the maternal DNA or RNA fraction in the mixed sample. In some embodiments, the method includes, for each of the plurality of loci, calculating (by computer, etc.) the expected amount of one or more alleles for that locus in the fetal DNA or RNA inherited from the mother in the mixed sample, using the identity of the allele present at that locus in the first or second homologous chromosome segment of maternal origin defined by the hypothesis as having been inherited by the fetus, the copy number of the first or second homologous chromosome segment of maternal origin defined by the hypothesis as having been inherited by the fetus, and a fraction of the fetal DNA or RNA in the mixed sample.
[0035] In some embodiments, the expected genetic data for each of the hypotheses includes the identity and quantity of one or more alleles present at each of multiple loci in the maternal DNA or RNA and fetal DNA or RNA in the mixed sample. In some embodiments, the method includes calculating the expected genetic data by determining the fetal DNA or RNA fraction and the maternal DNA or RNA fraction of the mixed sample. In some embodiments, the method includes calculating (by computer, etc.) the expected quantity of one or more alleles for that locus of maternal DNA or RNA in the mixed sample, using the identity of the allele present at that locus in the obtained maternal phase-determined genetic data and the maternal DNA or RNA fraction in the mixed sample for each of the multiple loci. In some embodiments, the method includes, for each of the multiple loci among the hypothesis, calculating the expected amount (by computer, etc.) of one or more alleles for that locus in the fetal DNA or RNA inherited from the mother in the mixed sample, using the identity of an allele present in the first or second homologous chromosome segment of maternal origin, which is hypothesized to have been inherited by the fetus; the copy number of the first or second homologous chromosome segment of maternal origin, which is hypothesized to have been inherited by the fetus; the identity of one or more alleles that may be present in that locus in the first or second homologous chromosome segment of paternal origin, which is hypothesized to have been inherited by the fetus; the copy number of the first or second homologous chromosome segment of paternal origin, which is hypothesized to have been inherited by the fetus; and fractions of fetal DNA or RNA in the mixed sample. In some embodiments, population frequencies are used to predict the identity of alleles in the first or second homologous chromosome segment of paternal origin. In some embodiments, the probability of each allele present at each locus of the first or second homologous chromosome segment of the paternal origin is considered to be equal.
[0036] In some embodiments, the method includes obtaining phase-determined genetic data for both the mother and father of the fetus. In some embodiments, the method includes listing a set of one or more hypotheses that define the copy number of a first maternal homologous chromosome segment or portion thereof in the fetal genome, the copy number of a second maternal homologous chromosome segment in the fetal genome, the copy number of a first paternal homologous chromosome segment or portion thereof in the fetal genome, the copy number of a second paternal homologous chromosome segment or portion thereof in the fetal genome, and the total copy number of the target chromosome segment present in the fetal genome. In some embodiments, the method includes, for each of the hypotheses, calculating (by computer, etc.) expected genetic data for multiple loci in the mixed sample from the obtained maternal and paternal phase-determined genetic data.
[0037] In some embodiments, for each of the hypotheses, the expected genetic data includes the identity and quantity of one or more alleles present at each of multiple loci in the maternal DNA or RNA and fetal DNA or RNA in the mixed sample. In some embodiments, the method includes calculating the expected genetic data by determining the DNA or RNA fraction of the mixed sample and the maternal DNA or RNA fraction of the fetus. In some embodiments, the method includes, for each of the multiple loci, calculating (by computer, etc.) the expected quantity of one or more alleles for that locus in the maternal DNA or RNA in the mixed sample, using the identity of the allele present at that locus in the obtained maternal phase-determined genetic data and the maternal DNA or RNA fraction in the mixed sample. In some embodiments, the method includes, for each of the multiple loci, calculating (by computer, etc.) the expected amount of one or more alleles for each locus of fetal DNA or RNA in the mixed sample, using the following: the identity of the allele of the first or second homologous chromosome segment of maternal origin that is hypothesized to have been inherited by the fetus; the copy number of the first or second homologous chromosome segment of maternal origin that is hypothesized to have been inherited by the fetus; the identity of the allele present at that locus of the first or second homologous chromosome segment of paternal origin that is hypothesized to have been inherited by the fetus; the copy number of the first or second homologous chromosome segment of paternal origin that is hypothesized to have been inherited by the fetus; and fractions of fetal DNA or RNA in the mixed sample.
[0038] In some embodiments, the method includes (using a computer, etc.) calculating the probability distribution of expected gene data for a plurality of loci in the mixed sample from the obtained phase-determined gene data from the (both) parents for each of the hypotheses. In some embodiments, the method includes increasing the probability in the probability distribution of a particular allele if that particular allele present at a first locus in the mixed sample is present at a first homologous segment of the parent and an allele present at a neighboring locus of the first homologous segment of the parent is observed in the obtained gene data of the mixed sample, or decreasing the probability in the probability distribution of a particular allele if that particular allele present at a first locus in the mixed sample is present at a first homologous segment of the parent and an allele present at a neighboring locus of the first homologous segment of the parent is not observed in the obtained gene data of the mixed sample. In some embodiments, the method includes increasing the probability of the probability distribution of a particular allele when a particular allele located at a second locus of the mixed sample is present in a second homologous segment of the parent and an allele located at a neighboring locus of the second homologous segment of the parent is observed in the genetic data of the obtained mixed sample, or decreasing the probability of the probability distribution of a particular allele when a particular allele located at a second locus of the mixed sample is present in a second homologous segment of the parent and an allele located at a neighboring locus of the second homologous segment of the parent is not observed in the genetic data of the obtained mixed sample.
[0039] In some embodiments, the method includes obtaining phase-determined genetic data from both the mother and father of the fetus. In some embodiments, the method includes listing a set of one or more hypotheses that define the copy number of a first homologous chromosome segment or portion thereof from the mother in the fetal genome, the copy number of a second homologous chromosome segment or portion thereof from the mother in the fetal genome, the copy number of a first homologous chromosome segment or portion thereof from the father in the fetal genome, the copy number of a second homologous chromosome segment or portion thereof from the father in the fetal genome, and the total copy number of the target chromosome segment present in the fetal genome. In some embodiments, the method includes, for each of the hypotheses, calculating (by computer, etc.) the probability distribution of expected genetic data for multiple loci of the mixed sample from the phase-determined genetic data obtained from the mother and father. In some embodiments, the method includes increasing the probability of the probability distribution of a particular allele if that particular allele located at a first locus in the mixed sample is present in the first homologous segment of the mother or father, and an allele located at a neighboring locus of the first homologous segment of that parent is observed in the genetic data of the obtained mixed sample, or decreasing the probability of the probability distribution of a particular allele if that particular allele located at a first locus in the mixed sample is present in the first homologous segment of the mother or father, and an allele located at a neighboring locus of the first homologous segment of that parent is not observed in the genetic data of the obtained mixed sample. In some embodiments, the method includes increasing the probability of the probability distribution of a particular allele when that particular allele located at a second locus of the mixed sample is present in the second homologous segment of the mother or father, and an allele located at a neighboring locus of the second homologous segment of that parent is observed in the genetic data of the obtained mixed sample, or decreasing the probability of the probability distribution of a particular allele when that particular allele located at a second locus of the mixed sample is present in the second homologous segment of the mother or father, and an allele located at a neighboring locus of the second homologous segment of that parent is not observed in the genetic data of the obtained mixed sample.
[0040] In some embodiments, the first seat and the seat adjacent to the first seat are co-separated. In some embodiments, the second seat and the seat adjacent to the second seat are co-separated. In some embodiments, no crossing is predicted between the first seat and the seat adjacent to the first seat. In some embodiments, no crossing is predicted between the second seat and the seat adjacent to the second seat. In some embodiments, the distance between the first seat and the seat adjacent to the first seat is less than 5mb, less than 1mb, less than 100kb, less than 10kb, less than 1kb, less than 0.1kb, or less than 0.01kb. In some embodiments, the distance between the second seat and the seat adjacent to the second seat is less than 5mb, less than 1mb, less than 100kb, less than 10kb, less than 1kb, less than 0.1kb, or less than 0.01kb.
[0041] In some embodiments, one or more crossovers occur during the formation of gametes that bring a copy of the target chromosome segment to the fetus, resulting in the target chromosome segment in the fetal genome, which includes a first homologous segment portion and a second homologous segment portion derived from the parent. In some embodiments, the set of hypotheses includes one or more hypotheses that define the copy number of the target chromosome segment in the fetal genome, which includes the first homologous segment portion and the second homologous segment portion derived from the parent.
[0042] In some embodiments, the expected genetic data of the mixed sample includes, for each of the hypotheses, the expected amounts of one or more alleles present at each of the multiple loci of the mixed sample.
[0043] In one aspect, the present invention is characterized by a method for determining, using phase-determined gene data, whether there is an over-recurrence of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of a given individual (for example, one or more cells, cell-free DNA, cell-free RNA, an individual suspected of having cancer, a fetus, or the genome of an embryo). In some embodiments, the method includes, simultaneously or in any order, (i) obtaining phase-determined genetic data for the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the first homologous chromosome segment, for each locus in the set of polymorphic loci of the first homologous chromosome segment; (ii) obtaining phase-determined genetic data for the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the second homologous chromosome segment, for each locus in the set of polymorphic loci of the second homologous chromosome segment; and (iii) obtaining measured allele genetic data, including the amount of each allele present at each locus in the set of polymorphic loci, in a sample of DNA or RNA derived from one or more cells from the individual or a mixed sample of cell-free DNA or RNA derived from two or more genetically distinct cells of the individual. In some embodiments, the method includes calculating allele ratios for one or more loci in the set of polymorphic loci that are heterozygous in at least one cell from which the sample originates. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing the measurement of one allele by the total measurement of all alleles at that locus. In some embodiments, the method includes determining whether there is an over-recurrence of the first homologous chromosome segment by comparing one or more calculated allele ratios for a given locus with the expected allele ratio (for example, the ratio expected for that locus if the first and second homologous chromosome segments are present in equal proportions). In some embodiments, the expected ratio is 0.5 for two-allelic loci.
[0044] In some embodiments of prenatal testing, the method includes, simultaneously or in any order, (i) obtaining phase-determined genetic data for a first homologous chromosome segment of the fetus's genome, including the identity of alleles present at each locus in the set of polymorphic loci in the first homologous chromosome segment, for each locus in the set of polymorphic loci in the first homologous chromosome segment; (ii) obtaining phase-determined genetic data for a second homologous chromosome segment of the fetus's genome, including the identity of alleles present at each locus in the set of polymorphic loci in the second homologous chromosome segment; and (iii) obtaining measured allele genetic data for a mixed sample of the fetus's maternal DNA or RNA, including fetal DNA or RNA and maternal DNA or RNA (for example, a mixed sample of cell-free DNA or RNA from a maternal blood sample, including fetal cell-free DNA or RNA and maternal cell-free DNA or RNA), including the amount of each allele present at each locus in the set of polymorphic loci. In some embodiments, the method includes calculating allele ratios for one or more loci in a set of polymorphic loci that are heterozygous in the fetus and / or heterozygous in the mother. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing a single measurement of the allele by the total measurement of all alleles at the locus. In some embodiments, the method includes determining whether there is an over-recurrence of the first homologous chromosome segment by comparing one or more calculated allele ratios for a given locus with the expected allele ratio (for example, the ratio expected for that locus if the first and second homologous chromosome segments are present in equal proportions).
[0045] In some embodiments, the calculated allele ratio indicates copy number overpopulation of the first homologous chromosome segment if (i) the allele ratio of the measured amount of alleles present at a locus on the first homologous chromosome, divided by the total measured amount of all alleles at that locus, is greater than the expected allele ratio for that locus, or (ii) the allele ratio of the measured amount of alleles present at a locus on the second homologous chromosome, divided by the total measured amount of all alleles at that locus, is less than the expected allele ratio for that locus. In some embodiments, the calculated allele ratio indicates that copy number overpopulation of the first homologous chromosome segment is not occurring if (i) the allele ratio of the measured amount of alleles present at a locus on the first homologous chromosome, divided by the total measured amount of all alleles at that locus, is less than or equal to the expected allele ratio for that locus, or (ii) the allele ratio of the measured amount of alleles present at a locus on the second homologous chromosome, divided by the total measured amount of all alleles at that locus, is greater than or equal to the expected allele ratio for that locus.
[0046] In some embodiments, determining whether there is a copy number overpopulation of the first homologous chromosome segment involves listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment. In some embodiments, the predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., the locus that is heterozygous in the fetus and / or the locus that is heterozygous in the mother) is estimated assuming that the degree of overpopulation is defined by each hypothesis. In some embodiments, the likelihood of a hypothesis being correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the most likely hypothesis is selected. In some embodiments, the expected distribution of the test statistic is calculated using the predicted allele ratio for each hypothesis. In some embodiments, the likelihood of a hypothesis being correct is calculated by comparing the test statistic calculated using the calculated allele ratio with the expected distribution of the test statistic calculated using the predicted allele ratio, and the most likely hypothesis is selected. In some embodiments, the predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., the locus that is heterozygous in the fetus and / or the locus that is heterozygous in the mother) is estimated assuming that the phase-determined gene data of the first homologous chromosome segment, the phase-determined gene data of the second homologous chromosome segment, and the excess frequency are defined by the hypothesis. In some embodiments, the likelihood that the hypothesis is correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the most likely hypothesis is selected.
[0047] In some embodiments, the ratio of DNA (or RNA) from one or more target cells to the total DNA (or RNA) in the sample is calculated. A typical ratio is the ratio of fetal DNA (or RNA) to the total DNA (or RNA) in the sample. In some embodiments, the ratio of fetal DNA to the total DNA in the sample is determined by measuring the amount of alleles at one or more loci where the fetus has an allele but the mother does not. In some embodiments, the ratio of fetal DNA to the total DNA in the sample is determined by measuring the difference in methylation between one or more maternal alleles and fetal alleles. In some embodiments, a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment is enumerated. In some embodiments, the predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother) is estimated assuming a calculated ratio of DNA or RNA, and the degree of overpopulation defined by the hypothesis is estimated for each hypothesis. In some embodiments, the likelihood of a hypothesis being correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the most likely hypothesis is selected. In some embodiments, the expected distribution of the test statistic calculated using the predicted allele ratio and the calculated ratio of DNA or RNA is estimated for each hypothesis. In some embodiments, the likelihood of a hypothesis being correct is determined by comparing the test statistic calculated using the calculated allele ratio and the calculated ratio of DNA or RNA with the expected distribution of the test statistic calculated using the predicted allele ratio and the calculated ratio of DNA or RNA, and the most likely hypothesis is selected.
[0048] In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment. In some embodiments, for each hypothesis, the method includes estimating either an expected distribution of a test statistic calculated using the predicted allele ratio and the possible ratio of DNA or RNA from one or more target cells (e.g., fetal cells) to the total DNA or RNA in the sample, for (i) a predicted allele ratio of a locus that is heterozygous in at least one cell assuming the degree of overpopulation is defined by the hypothesis (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother), or (ii) a possible ratio of one or more DNA or RNA (e.g., a ratio of fetal DNA or RNA to the total DNA or RNA in the sample). In some embodiments, data fit is calculated by comparing either (i) the calculated allele ratios to the predicted allele ratios or (ii) a test statistic calculated using the calculated allele ratios and the possible ratios of the DNA or RNA with the expected distribution of the test statistic calculated using the predicted allele ratios and the possible ratios of the DNA or RNA. In some embodiments, one or more hypotheses are ranked by the data fit, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm (e.g., a search algorithm) is used in one or more of the steps of calculating the data fit, ranking the hypotheses, and selecting the highest-ranked hypothesis. In some embodiments, the data fit is a fit to a beta-binomial distribution or a fit to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, maximum posterior probability estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained gene data and the expected gene data.
[0049] In some embodiments, the method includes creating a range of possible ratios (e.g., the ratio of fetal DNA or RNA to the total DNA or RNA in the sample) from a lower limit to an upper limit for the ratio of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample. In some embodiments, a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment is enumerated. In some embodiments, the method includes, for each possible ratio of DNA or RNA in the range and each hypothesis, estimating either (i) a predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother) given the possible ratio of DNA or RNA and the degree of overpopulation defined by the hypothesis, or (ii) an expected distribution of a test statistic calculated using the predicted allele ratio and the possible ratio of DNA or RNA. In some embodiments, the method includes calculating the likelihood that a hypothesis is correct for each ratio of the possible DNA or RNA in a given category and each hypothesis by comparing (i) a calculated allele ratio to the predicted allele ratio or (ii) a test statistic calculated using the calculated allele ratio and the possible ratio of the DNA or RNA with the expected distribution of the test statistic calculated using the predicted allele ratio and the possible ratio of the DNA or RNA. In some embodiments, the composite probability of each hypothesis is determined by combining the probabilities of the hypothesis for each ratio of the possible ratios in the given category, and the hypothesis with the highest composite probability is selected. In some embodiments, the composite probability of each hypothesis is determined by taking a weighted average of the hypothesis probabilities for a particular possible ratio based on the likelihood that the possible ratio is the correct ratio.
[0050] In one aspect, the present invention features a method for determining the copy number of chromosomes or chromosome segments in the genome of one or more cells from a given individual using phase-determined or unphased genetic data. In some embodiments, the method includes obtaining genetic data of a set of polymorphic loci in the chromosome or chromosome segment in a sample by measuring the amount of each allele present at each locus. In some embodiments, the sample is a mixed sample of cell-free DNA from the individual, including a sample of DNA or RNA from one or more cells from the individual, or cell-free DNA from two or more genetically distinct cells. In some embodiments, an allele ratio is calculated for heterozygous loci in at least one cell from which the sample originates. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing the measured amount of one allele by the total measured amount of all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one metric of an allele (e.g., an allele of the first homologous chromosome segment) by the metric of one or more other alleles of that locus (e.g., an allele of the second homologous chromosome segment). In some embodiments, a set of one or more hypotheses defining the copy number of chromosomes or chromosome segments of one or more genomes of the cell is enumerated. In some embodiments, the most likely hypothesis is selected based on the test statistic, thereby determining the copy number of chromosomes or chromosome segments of one or more genomes of the cell.
[0051] In one aspect, the present invention is characterized by a method for determining the copy number of a chromosome or chromosome segment of a fetal genome (e.g., a fetus in the womb of a pregnant mother) using phase-determined or phase-undetermined genetic data. In some embodiments, the method relates to obtaining genetic data of a set of polymorphic loci in the chromosome or chromosome segment in a sample by measuring the amount of each allele present at each locus. In some embodiments, the sample is a mixed DNA sample containing fetal DNA or RNA and maternal DNA or RNA derived from the mother of the fetus (e.g., a mixed sample of cell-free DNA or RNA containing fetal cell-free DNA or RNA and maternal cell-free DNA or RNA derived from a maternal blood sample). In some embodiments, allele ratios are calculated for the heterozygous loci of the fetus and / or the heterozygous loci of the mother. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing the measured amount of one allele by the total measured amount of all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele (e.g., an allele in the first homologous chromosome segment) by the measurements of one or more other alleles in that locus (e.g., an allele in the second homologous chromosome segment). In some embodiments, a set of one or more hypotheses defining the copy number of a chromosome or chromosome segment in the fetal genome is enumerated. In some embodiments, the most likely hypothesis is selected based on the test statistic, thereby determining the copy number of a chromosome or chromosome segment in the fetal genome.
[0052] In some embodiments, a hypothesis is selected if the probability that the test statistic belongs to the test statistic distribution of a hypothesis exceeds an upper threshold, and one or more hypotheses are rejected if the probability that the test statistic belongs to the test statistic distribution of one or more hypotheses is below a lower threshold. Alternatively, if the probability that the test statistic belongs to the test statistic distribution of a hypothesis is between the lower and upper thresholds, or if the probability cannot be determined with sufficiently high confidence, the hypothesis is neither selected nor rejected. In some embodiments, the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, the total measured amount of all alleles for one or more loci is compared to a reference amount to determine whether the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, the magnitude of the difference between the calculated allele ratio and the expected allele ratio for one or more loci is used to determine whether the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, it is determined that the first and second homologous chromosome segments are present in equal proportions if there is no copy number overpopulation of the first homologous chromosome segment and no overpopulation of the second homologous chromosome segment (e.g., in the genome of a cell, cell-free DNA, cell-free RNA, individual, fetus, or embryo).
[0053] In some embodiments, the ratio of DNA derived from one or more target cells to the total DNA in the sample is determined based on the total or relative amount of one or more alleles present at one or more loci where the target and non-target cells are predicted to be dichromosomes, and the genotype of the target cells differs from that of the non-target cells. In some embodiments, this ratio is used to determine whether the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, this ratio is used to determine the duplicated chromosome segment or the excess copy number of the chromosome. In some embodiments, the phase-determined gene data includes probabilistic data. In some embodiments, obtaining phase-determined genetic data for a first homologous chromosome segment and / or a second homologous chromosome segment in the fetal genome includes obtaining phase-determined genetic data for a first homologous chromosome segment and / or a second homologous chromosome segment in the genomes of one or both of the fetal biological parents and inferring which homologous chromosome segment the fetus inherited from which of its biological parents. In some embodiments, the probability of one or more crossovers (e.g., one, two, three, or four) that may have occurred during the formation of the gamete that resulted in a copy of the first or second homologous chromosome segment in the fetal individual is used to infer which homologous chromosome segment the fetus inherited from which of its biological parents. In some embodiments, phased genetic data of the maternal and / or paternal fetus can be obtained using techniques selected from a group consisting of digital PCR, haplotype inference using population-based haplotype frequencies, haplotype determination using haploid cells (e.g., sperm or oocytes), haplotype determination using genetic data from one or more first-degree relatives, and combinations thereof. In some embodiments, phased genetic data of the individual can be obtained by phase-determining some or all of the regions corresponding to deletions or duplications in a sample derived from the individual. In some embodiments, phased genetic data of the fetus can be obtained by phase-determining some or all of the regions corresponding to deletions or duplications in a sample derived from the fetus or the maternal fetus.In some embodiments, obtaining phase-determined genetic data for the first and second homologous chromosome segments involves inferring the identity of alleles present in one of these chromosome segments and the identity of alleles present in the other chromosome segment. In some embodiments, alleles not present in the first homologous chromosome segment, derived from undetermined genetic data, are assigned to the second homologous chromosome segment. For example, if an individual's genotype is (AB,AB) and the phase-determined data for that individual indicates that the first haplotype is (A,A), then the other haplotype can be inferred to be (B,B). In some embodiments, if only one allele is measured at a locus, that allele is determined to be part of both the first and second homologous chromosome segments (for example, if the genotype is AA at a locus, both haplotypes have allele A). In some embodiments, the phase-determined genetic data of the individual includes determining, for example, the sequences of recombination-prone regions and optionally the sequences of regions adjacent to recombination-prone regions to determine whether one or more possible chromosomal crossovers have occurred. In some embodiments, recombination events are detected using one of the primer libraries of the present invention to determine which haplotype blocks are present in the genome of a given individual.
[0054] In some embodiments, the method includes using a joint distribution model (e.g., a joint distribution model that considers linkage between loci), performing linkage analysis, using a binomial distribution model, using a beta-binomial distribution model, and / or using the likelihood of crossover that occurred during meiosis that produced gametes that formed the embryo (e.g., modeling the degree of dependence between polymorphic alleles in a target chromosome or chromosome segment using the probability of chromosomes crossing over at different locations on the chromosome).
[0055] In some embodiments, one or more calculated allele ratios for cell-free DNA or cell-free RNA represent the allele ratio of the corresponding DNA or RNA in the cell from which the cell-free DNA or cell-free RNA originates. In some embodiments, one or more calculated allele ratios for cell-free DNA or cell-free RNA represent the corresponding allele ratio in the genome of the individual. In some embodiments, if the measured genetic data indicates that two or more different alleles are present at that locus in the sample (e.g., in a cell-free DNA or cell-free RNA sample), the allele ratio is calculated only or compared to an expected allele ratio. In some embodiments, if the locus is heterozygous in at least one of the cells from which the sample originates (e.g., a locus that is heterozygous in the fetus and / or in its mother), the allele ratio is calculated only or compared to an expected allele ratio. In some embodiments, if the locus is heterozygous in the fetus, the allele ratio is calculated only or compared to an expected allele ratio. In some embodiments, allele ratios are calculated and compared to the expected allele ratios for homozygous loci. For example, the allele ratios of loci predicted to be homozygous for a particular individual being tested (or both the fetus and the pregnant mother) may be analyzed to determine the amount of noise or error in the system.
[0056] In some embodiments, at least 10, 50, 100, 200, 300, 500, 750, 1,000, 2,000, 3,000, or 4,000 or more (e.g., SNP) loci are analyzed for the chromosome or chromosome segment of interest. In some embodiments, the average number of (e.g., SNP) loci per megabase pair (mb) of the chromosome or chromosome segment of interest is at least 1, 10, 25, 50, 100, 150, 200, 300, 500, 750, or 1,000 or more per megabase pair. In some embodiments, the average number of loci (e.g., SNPs) per megabase pair of the chromosome or chromosomal segment of interest is between 1 and 500 per megabase pair (e.g., between 1 and 50, between 50 and 100, between 100 and 200, between 200 and 400, between 200 and 300, or between 300 and 400, including the values at each end). In some embodiments, loci of multiple possible deletions or duplications are analyzed to increase the sensitivity and / or specificity of CNV determination compared to analyzing only one locus or only a few adjacent loci. In some embodiments, only the two most common alleles present at each locus are measured, or only these alleles are used to determine the calculated allele ratio. In some embodiments, locus amplification is performed using a polymerase having 5'→3' exonuclease activity and / or low-chain substitution activity (e.g., DNA polymerase, RNA polymerase, or reverse transcriptase). In some embodiments, the genetic data of the measured allele is obtained by (i) sequencing the DNA or RNA in the sample, (ii) amplifying the DNA or RNA in the sample and sequencing the amplified DNA, or (iii) amplifying the DNA or RNA in the sample, ligating the PCR product, and sequencing the ligated product.In some embodiments, alleles of polymorphic loci (e.g., SNPs) are identified using one or more of the following methods: sequencing (e.g., nanopore sequencing, Halcyon molecular sequencing), SNP arrays, real-time PCR, TaqMan, Nanostring nCounter® analyzer, Illumina GoldenGate Genotyping Assay using differential DNA polymerase and ligase, ligation-mediated PCR, and ligated reverse probes (sometimes called LIPs, pre-circularized probes, pre-circularized probes, circularized probes, padlock probes, or molecular inversion probes (MIPs)). In some embodiments, two or more (e.g., three or four) target amplification products are ligated together, and the ligate product is sequenced. In some embodiments, measurements of different alleles at the same locus are adjusted for differences in metabolism, apoptosis, histones, inactivity, and / or amplification between these alleles (e.g., differences in amplification efficiency between different alleles at the same locus). In some embodiments, this adjustment is performed before calculating the allele ratio of the obtained gene data or before comparing the measured gene data with expected gene data.
[0057] In some embodiments, the method also includes determining the presence or absence of one or more risk factors for a disease or disorder. In some embodiments, the method also includes determining the presence or absence of one or more polymorphisms or mutations associated with a disease or disorder or an increased risk of such disease or disorder. In some embodiments, the method also includes determining the total amount of cell-free DNA, cell-free mitochondrial DNA, cell-free nuclear DNA, cell-free RNA, miRNA, or any combination thereof. In some embodiments, the method includes determining the amount of one or more cell-free DNA, cell-free mitochondrial DNA, cell-free nuclear DNA, cell-free RNA, and / or miRNA molecules of interest (e.g., molecules having polymorphisms or mutations associated with a disease or disorder or an increased risk of such disease or disorder). In some embodiments, a tumor DNA fraction in total DNA (e.g., a tumor cell-free DNA fraction in total cell-free DNA, a tumor cell-free DNA fraction with a specific mutation in total cell-free DNA, etc.) is determined. In some embodiments, this tumor fraction is used to determine the stage of cancer (since a higher tumor fraction can be associated with a more advanced stage of cancer). In some embodiments, the method also includes determining the total amount of DNA or RNA. In some embodiments, the method includes determining the amount of one or more methylated DNA or RNA molecules of interest (e.g., molecules having polymorphisms or mutations associated with a disease or disorder or an increased risk of such disease or disorder). In some embodiments, the method includes determining whether or not there is an alteration in the integrity of the DNA. In some embodiments, the method also includes determining the total amount of mRNA splicing. In some embodiments, the method includes determining the amount of mRNA splicing or detecting another mRNA splicing corresponding to one or more RNA molecules of interest (e.g., molecules having polymorphisms or mutations associated with a disease or disorder or an increased risk of such disease or disorder).
[0058] In some embodiments, the present invention provides a method for detecting a cancerous phenotype in an individual, characterized in that the cancerous phenotype is defined by the presence of at least one mutation from a set of mutations. In some embodiments, the method comprises obtaining a DNA or RNA measurement on a sample of DNA or RNA derived from one or more cells from the individual when one or more of the cells are suspected to have a cancerous phenotype, and analyzing the DNA or RNA measurement to determine the likelihood that at least one of the cells has each mutation from the set of mutations. In some embodiments, the method comprises determining that the individual has the cancerous phenotype if (i) for at least one of the mutations, the likelihood that at least one of the cells has that mutation is greater than a threshold, or (ii) for at least one of the mutations, the likelihood that at least one of the cells has that mutation is less than the threshold, and for multiple mutations, the combination of likelihoods that at least one of the cells has at least one of the mutations is greater than the threshold. In some embodiments, one or more cells have a subset or all of the mutations within the set of mutations. In some embodiments, the subset of the mutations is associated with cancer or an increased risk of cancer. In some embodiments, the sample includes cell-free DNA or RNA. In some embodiments, the DNA or RNA measurement includes a measurement of a set of polymorphic loci in one or more chromosomes or chromosomal segments (e.g., the amount of each allele present in each locus).
[0059] In one aspect, the present invention features a method for selecting a treatment for the treatment, stabilization, or prevention of a disease or disorder in a mammal. In some embodiments, the method includes determining whether there is an over-recurrence of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment using any of the methods disclosed herein. In some embodiments, a treatment is selected for the mammal (e.g., a treatment for a disease or disorder associated with the over-recurrence of the first homologous chromosome segment).
[0060] In one aspect, the present invention features methods for preventing, delaying, stabilizing, or treating diseases or disorders in mammals. In some embodiments, the method includes determining, using any of the methods disclosed herein, whether there is an over-recurrence of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment. In some embodiments, a treatment (e.g., a treatment for a disease or disorder associated with the over-recurrence of the first homologous chromosome segment) is selected and administered to the mammal.
[0061] In some embodiments, treating, stabilizing, or preventing a disease or disorder includes preventing or delaying the initial onset or subsequent events of the disease or disorder, increasing disease-free survival time from the disappearance of the condition to its recurrence, stabilizing or reducing adverse symptoms associated with the condition, or inhibiting or stabilizing the progression of the condition. In some embodiments, at least 20, 40, 60, 80, 90, or 95% of treated subjects are completely remission, with all signs of the condition gone. In some embodiments, the survival time of subjects after being diagnosed with and treated for a condition is extended by at least 20, 40, 60, 80, 100, 200, or 500% of the mean survival time of untreated subjects or (ii) the mean survival time of subjects treated with other therapies.
[0062] In some embodiments, treating, stabilizing, or preventing cancer includes reducing or stabilizing the size of a tumor (e.g., a benign or malignant tumor), delaying or preventing tumor growth, reducing or stabilizing the number of tumor cells, increasing disease-free survival time from tumor disappearance to recurrence, preventing the initial onset or subsequent onset of a tumor, or reducing or stabilizing tumor-associated adverse conditions. In one embodiment, the number of cancerous cells remaining after treatment is reduced by at least 10, 20, 40, 60, 80, or 100% compared to the initial number of cancerous cells, as measured using any standard assay. In some embodiments, the number of cancerous cells induced by the treatment of the present invention is reduced by at least 2, 5, 10, 20, or more than 50 times compared to the number of non-cancer cells. In some embodiments, the number of cancerous cells present after treatment is reduced by at least 2, 5, 10, 20, or 50 times compared to the number of cancerous cells present after administration of a control (e.g., saline or buffer). In some embodiments, the method of the present invention resulted in a 10, 20, 40, 60, 80, or 100% reduction in tumor size as determined by standard methods. In some embodiments, at least 10, 20, 40, 60, 80, 90, or 95% of treated subjects showed complete sedation with no detectable cancer cells. In some embodiments, the cancer was completely recurrence-free, or remained recurrence-free for at least 2, 5, 10, 15, or 20 years. In some embodiments, the survival time of subjects diagnosed with cancer and treated with the treatment method of the present invention was extended by at least 10, 20, 40, 60, 80, 100, 200, or 500% compared to (i) the mean survival time of untreated subjects or (ii) the mean survival time of subjects treated with other treatments.
[0063] In one aspect, the present invention features a method for stratifying subjects to be included in a clinical trial for the treatment, stabilization, or prevention of a mammalian disease or disorder. In some embodiments, the method includes determining, using any of the methods disclosed herein, whether there is an over-reproduction of a first homologous chromosome segment compared to a second homologous chromosome segment before, during, or after the clinical trial. In some embodiments, the subjects are assigned to subgroups for the clinical trial based on the presence or absence of an over-reproduction of the first homologous chromosome segment in their genome.
[0064] In some embodiments, the disease or disorder is selected from the group consisting of cancer, mental disorders, learning disabilities (e.g., idiopathic learning disabilities), intellectual disability, developmental delay, autism, neurodegenerative diseases or disorders, schizophrenia, physical disabilities, autoimmune diseases or disorders, systemic lupus erythematosus, psoriasis, Crohn's disease, glomerulonephritis, HIV infection, AIDS, and combinations thereof. In some aspects, the disease or disorder is DiGeorge syndrome, DiGeorge II syndrome, DiGeorge / Velopal-Cardios-Face (VCFS) syndrome, Prader-Willi syndrome, Angelman syndrome, Beckwith-Wiedemann syndrome, 1p36 deletion syndrome, 2q37 deletion syndrome, 3q29 deletion syndrome, 9q34 deletion syndrome, 17q21 / 31 deletion syndrome, Cri-du-chat syndrome, Jacobsen syndrome, Miller-Dicker syndrome The group consists of Dieker syndrome, Phelan-McDermid syndrome, Smith-Magenis syndrome, Tumor-Aniridia-Urogenitourinary Malformation-Intellectual Retardation (WAGR) syndrome, Wolf-Hirschhorn syndrome, Williams syndrome, Williams-Beuren syndrome, Miller-Dieker syndrome, Phelan-McDermid syndrome, Smith-Magenis syndrome, Down syndrome, Edward syndrome, Patau syndrome, Klinefelter syndrome, Turner syndrome, 47,XXX syndrome, 47,XYY syndrome, Sotos syndrome, and combinations thereof.In some embodiments, the method determines the presence or absence of one or more of the following chromosomal abnormalities: zero chromosome, monochromosome, uniparental dichromosome, trichromosome, congruent trichromosome, partially congruent trichromosome, maternal trichromosome, parental trichromosome, triploid, mosaic tetrachromosome, congruent tetrachromosome, partially congruent tetrachromosome, other aneuploidy, unbalanced translocation, balanced translocation, insertion, deletion, recombination, and combinations thereof. In some embodiments, the chromosomal abnormality is any deviation from the most common copy number of a particular chromosome or chromosome segment, for example, in human somatic cells, any deviation from 2 copies can be considered a chromosomal abnormality. In some embodiments, the method determines the presence or absence of euploidy. In some embodiments, the copy number hypothesis includes one or more copy number hypotheses for singleton pregnancies. In some embodiments, the copy number hypothesis includes one or more copy number hypotheses for multiple pregnancies, such as twin pregnancies (e.g., identical twins, sibling twins, disappearing twins). In some embodiments, the copy number hypothesis includes all fetuses that are euploid in a multiple pregnancy, all fetuses that are aneuploid in a multiple pregnancy (e.g., any aneuploidy disclosed herein), and / or one or more fetuses that are euploid in a multiple pregnancy, and one or more fetuses that are aneuploid in a multiple pregnancy. In some embodiments, the copy number hypothesis includes identical twins (also known as monozygous twins) or sibling twins (also known as dizygous twins). In some embodiments, the copy number hypothesis includes molar pregnancies (e.g., complete molar pregnancies, partial molar pregnancies, etc.). In some embodiments, the chromosome segment of interest is an entire chromosome. In some embodiments, the chromosome or chromosome segment is selected from the group consisting of chromosome 13, chromosome 18, chromosome 21, X chromosome, Y chromosome, segments thereof, and combinations thereof. In some embodiments, the first homologous chromosome segment and the second homologous chromosome segment are a pair of homologous chromosome segments containing the chromosome segment of interest. In some embodiments, the first homologous chromosome segment and the second homologous chromosome segment are a pair of homologous chromosomes of interest. In some embodiments, confidence is calculated for the determination of CNV or for the diagnosis of disease or disorder.
[0065] In some embodiments, the deletion is at least 0.01kb, 0.1kb, 1kb, 10kb, 100kb, 1mb, 2mb, 3mb, 5mb, 10mb, 15mb, 20mb, 30mb, or 40mb. In some embodiments, the deletion is between 1kb and 40mb, for example, between 1kb and 100kb, 100kb and 1mb, between 1 and 5mb, 5 and 10mb, 10 and 15mb, 15 and 20mb, 20 and 25mb, 25 and 30mb, or 30 and 40mb (including numerical values). In some embodiments, one copy of the chromosome segment is deleted and one copy is present. In some embodiments, two copies of the chromosome segment are deleted. In some embodiments, an entire chromosome is deleted.
[0066] In some embodiments, the duplication is at least 0.01kb, 0.1kb, 1kb, 10kb, 100kb, 1mb, 2mb, 3mb, 5mb, 10mb, 15mb, 20mb, 30mb, or 40mb. In some embodiments, the duplication is between 1kb and 40mb, for example, 1kb and 100kb, 100kb and 1mb, 1 and 5mb, 5 and 10mb, 10 and 15mb, 15 and 20mb, 20 and 25mb, 25 and 30mb, or 30 and 40mb (including numerical values). In some embodiments, the duplication of the chromosome segment occurs once. In some embodiments, the duplication of the chromosome segment occurs two or more times (for example, two, three, four, or five times). In some embodiments, an entire chromosome is duplicated. In some embodiments, a region of the first homologous segment is deleted, and the same region or other regions of the second homologous segment overlap. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 98, 99, and 100% of the SNVs examined are conversion mutations and not translocation mutations.
[0067] In some embodiments, the sample comprises DNA and / or RNA derived from (i) one or more target cells or (ii) one or more non-target cells. In some embodiments, the sample is a mixed sample comprising DNA and / or RNA derived from one or more target cells and one or more non-target cells. In some embodiments, the target cells are cells having CNVs (e.g., deletion or duplication of the target), and the non-target cells are cells that do not have the copy number diversity of the target. In some embodiments where the one or more target cells are cancer cells and the one or more non-target cells are non-cancerous cells, the method comprises determining whether there is copy number overpopulation of the first homologous chromosome segment in the genome of one or more of the cancer cells. In some embodiments where the one or more target cells are genetically identical cancer cells and the one or more non-target cells are non-cancerous cells, the method comprises determining whether there is copy number overpopulation of the first homologous chromosome segment in the genome of the cancer cells. In some embodiments where the one or more target cells are genetically dissimilar cancer cells and the one or more non-target cells are non-cancer cells, the method includes determining whether there is a copy number overpopulation of the first homologous chromosome segment in the genome of one or more of the genetically dissimilar cancer cells. In some embodiments where the sample includes cell-free DNA derived from a mixture of one or more cancer cells and one or more non-cancer cells, the method includes determining whether there is a copy number overpopulation of the first homologous chromosome segment in the genome of one or more of the cancer cells. In some embodiments where the one or more target cells are genetically identical fetal cells and the one or more non-target cells are maternal cells, the method includes determining whether there is a copy number overpopulation of the first homologous chromosome segment in the genome of the fetal cells. In some embodiments, where the one or more target cells are genetically dissimilar fetal cells and the one or more non-target cells are maternal cells, the method includes determining whether there is a copy number overpopulation of the first homologous chromosome segment in the genome of one or more of the genetically dissimilar fetal cells.Since the cells of most individuals contain a nearly identical set of nuclear DNA, in some aspects the term “target cell” may be used interchangeably with the term “individual.” Cancer cells have a different genotype from the host individual. In this case, the cancer itself may be considered an individual. Also, many cancers are heterogeneous, meaning that each cell in a tumor is genetically different from other cells in the same tumor. In this case, genetically identical regions can be considered as distinct individuals. Alternatively, the cancer may be considered a single individual containing a mixture of cells with different genomes. Non-target cells are usually euploid, but not always.
[0068] In some embodiments, the sample is obtained from a whole blood sample or fraction thereof from the mother, cells isolated from a maternal blood sample, amniocentesis sample, product of a conception sample, placental tissue sample, chorionic villi sample, placental membrane sample, cervical mucus sample, or fetal-derived sample. In some embodiments, the sample includes cell-free DNA obtained from the maternal-derived blood sample or fraction thereof. In some embodiments, the sample includes nuclear DNA obtained from a mixture of fetal and maternal cells. In some embodiments, the sample is obtained from a fraction of maternal blood containing nucleated cells enriched with fetal cells. In some embodiments, the sample is divided into multiple fractions (e.g., two, three, four, five or more fractions) and each is analyzed using the method of the present invention. If each fraction shows the same result (e.g., presence or absence of one or more CNVs of interest), the reliability of the result is increased. If each fraction shows different results, the sample may be re-analyzed, or other samples may be taken and analyzed from the same subject.
[0069] Examples of subjects include mammals (e.g., humans and mammals for veterinary purposes). In some embodiments, the mammals are primates (e.g., humans, monkeys, gorillas, apes, lemurs, etc.), cattle, horses, pigs, dogs, or cats.
[0070] In some embodiments, any of the methods described above includes preparing a report (e.g., a written report or an electronic report) disclosing the results of the methods described above of the present invention (e.g., whether there are any omissions or duplicates).
[0071] In some embodiments, any of the methods described above include performing a clinical action based on the results of the methods described above of the present invention (e.g., whether or not there are deletions or duplications). In some embodiments, based on the results of the methods described above of the present invention, if an embryo or fetus has one or more polymorphisms or mutations of interest (e.g., CNVs), the clinical action includes performing additional tests (e.g., tests to confirm the presence of polymorphisms or mutations), but does not include transferring the embryo for in vitro fertilization, transferring another embryo for in vitro fertilization, terminating the pregnancy, preparing the child for special needs, or performing interventions designed to reduce the severity of the phenotypic presentation of a genetic disorder. In some embodiments, the clinical procedures are selected from a group consisting of ultrasound diagnosis, amniocentesis of a fetus, amniocentesis of a fetus that has inherited genetic material from the mother and / or father, chorionic villus biopsy of a fetus, chorionic villus biopsy of a fetus that has inherited genetic material from the mother and / or father, in vitro fertilization, preimplantation genetic diagnosis of one or more embryos that have inherited genetic material from the mother and / or father, maternal chromosome analysis, paternal chromosome analysis, fetal echocardiography (e.g., echocardiography of a fetus with trichromosomal 21, 18, or 13, monochromosomal X, or microdeletion), and combinations thereof.In some embodiments, the clinical practice may include administering growth hormone to a child born with monochromosome X (e.g., starting administration at approximately 9 months of age), administering calcium to a child born with 22q deletion (e.g., DiGeorge syndrome), administering androgens such as testosterone to a child born with 47,XXY (e.g., injecting 25 mg of testosterone enanthate per month into an infant or toddler for 3 months), performing cancer screenings on women with complete or partial molar pregnancies (e.g., triploid fetuses), performing cancer treatment (with chemotherapy drugs, etc.) on women with complete or partial molar pregnancies (e.g., triploid fetuses), and administering one or more X-related genetic disorders (e.g., Duchenne muscular dystrophy (DMD), adrenoleukodystrophy) to a fetus determined to be male (e.g., a fetus determined to be male using the method of the present invention). The following are selected from a group consisting of screening tests for X-related diseases (e.g., hemophilia), amniocentesis of male fetuses at risk of X-related diseases, administration of dexamethasone to women pregnant with female fetuses at risk of congenital adrenal hyperplasia (e.g., fetuses determined to be female using the method of the present invention), amniocentesis of male fetuses at risk of congenital adrenal hyperplasia, administration of dead vaccines (instead of live vaccines) or avoidance of certain vaccines to children born with immunodeficiency (or suspected immunodeficiency) due to 22q11.2 deletion, implementation of occupational therapy and / or physical therapy, early intervention in education, transfer of infants to a tertiary care center with neonatal intensive care units and / or pediatric specialists capable of handling deliveries, behavioral interventions for newborns (e.g., children with XXX, XXY, or XYY), and combinations thereof.
[0072] In some cases, women diagnosed with a multiple pregnancy (e.g., twins) undergo ultrasound or other screening tests to determine whether two or more fetuses are monochorionic. Identical twins develop from the ovulation and fertilization of a single oocyte, followed by cell division of the fertilized egg. Placental formation can be dichorionic or monochorionic. Fraternal twins develop from the ovulation and fertilization of two oocytes, but usually a dichorionic placenta is formed. Monochorionic twins are at risk of twin-to-twin transfusion syndrome. These syndromes can lead to an imbalance in blood distribution between the fetuses, resulting in differences in growth and development between the twins, and in some cases, stillbirth. Therefore, it is desirable that twins determined to be identical twins using the method of the present invention be examined (for example by ultrasound) to determine whether they are monochorionic twins, and if they are monochorionic twins, these twins can be monitored (for example by ultrasound every other week from the 16th week) for signs of twin-to-twin transfusion syndrome.
[0073] In some embodiments of the present invention, if the embryo or fetus does not have one or more polymorphisms or mutations of interest (e.g., CNVs) as a result of the method described above, the clinical procedure includes transferring the embryo for in vitro fertilization or continuing the pregnancy. In some embodiments, the clinical procedure is an additional examination to confirm the absence of the polymorphisms or mutations by performing an action selected from the group consisting of ultrasound, amniocentesis, chorionic villus biopsy, and combinations thereof.
[0074] In some embodiments, the clinical action includes performing additional tests or administering one or more treatments for the disease or disorder (e.g., treatments for cancer, treatments for a specific type of cancer or specific type of mutation to which the individual has been diagnosed, or any of the treatments disclosed herein). In some embodiments, the clinical action is an additional test to confirm the presence or absence of the polymorphism or mutation, selected from the group consisting of biopsy, surgery, medical imaging (e.g., mammography, ultrasound, etc.), and combinations thereof.
[0075] In some embodiments, the additional testing includes performing the same or different methods (e.g., any of the methods disclosed herein) to determine the presence or absence of polymorphisms or mutations (e.g., CNVs) (e.g., testing a second fraction of the same sample tested, or a different sample from the same individual (e.g., the same pregnant mother, fetus, embryo, or individual with an increased risk of cancer)). In some embodiments, the additional testing (e.g., additional testing to confirm the presence of a more likely polymorphism or mutation) is performed on individuals where the probability of polymorphisms or mutations (e.g., CNVs) exceeds a threshold. In some embodiments, the additional testing (e.g., additional testing to confirm the presence of a more likely polymorphism or mutation) is performed on individuals where the confidence or Z-value for determining polymorphisms or mutations (e.g., CNVs) exceeds a threshold. In some embodiments, the additional testing (e.g., additional testing to increase the confidence that the initial result is correct) is performed on individuals where the confidence or Z-value for determining polymorphisms or mutations (e.g., CNVs) is between a minimum and maximum threshold. In some embodiments, the additional tests are performed on individuals for whom the confidence level for determining the presence or absence of polymorphisms or mutations (e.g., CNVs) is below a threshold (e.g., "no call," which is the result of not being able to determine the presence or absence of CNVs with sufficient confidence). The Z-value is calculated, for example, using Chiu et al. BMJ 2011;342:c7401 (which is incorporated herein by reference in its entirety). Here, chromosome 21 is used as an example, but it may be substituted with any other chromosome or chromosome segment in the test sample. The Z-value for the percentage of chromosome 21 in the test case = ((Percentage of chromosome 21 in the test case) - (Mean value of the percentage of chromosome 21 in the control group)) / (Standard deviation of the percentage of chromosome 21 in the control group) In some embodiments, the additional testing is performed on individuals whose initial sample did not meet quality control guidelines or contained a fetal or tumor fraction below a threshold. In some embodiments, the method includes selecting individuals to undergo additional testing based on the results of the method of the present invention, the probability of the results, the confidence level of the results, or the Z-value, and performing the additional testing on those individuals (e.g., the same or a different sample). In some embodiments, individuals diagnosed with a disease or disorder (e.g., cancer) are subjected to repeated testing at multiple time points using the method of the present invention or well-known testing methods for testing for the disease or disorder to monitor for progression, remission, or recurrence of the disease or disorder.
[0076] In one aspect, the present invention is characterized by a report (e.g., a written report or an electronic report) disclosing the results of the method of the present invention (e.g., whether there are any omissions or duplicates).
[0077] In various embodiments, the primer extension reaction, i.e., the polymerase chain reaction, involves adding one or more nucleotides with a polymerase. In some embodiments, the primers are in solution. In some embodiments, the primers are in solution and are not immobilized on a solid support. In some embodiments, the primers are not part of a microarray. In various embodiments, the primer extension reaction, i.e., the polymerase chain reaction, does not involve ligation-mediated PCR. In various embodiments, the primer extension reaction, i.e., the polymerase chain reaction, does not involve ligating two primers with a ligase. In various embodiments, the primers do not include a ligated reverse probe (sometimes called a LIP, cyclized probe, pre-cyclized probe, cyclized probe, padlock probe, or molecular inversion probe (MIP)).
[0078] The aspects and embodiments of the present invention described herein include any combination of two or more aspects or embodiments of the present invention. definition
[0079] A single nucleotide polymorphism (SNP) refers to a single nucleotide that may differ in the genomes of two individuals of the same species. The use of this term is not intended to imply any limitation on the frequency with which each variant occurs.
[0080] "Sequence" refers to a DNA sequence or gene sequence. This term may also refer to the primary physical structure of an individual's DNA molecule or DNA strand. This term may also refer to the nucleotide sequence found in a DNA molecule, or the complementary strand of that DNA molecule. This term may also refer to the information contained within that DNA molecule as a computer representation.
[0081] A "locus" refers to a specific region of an individual's DNA, which may include a SNP, a site of potential insertion or deletion, or any other site of related genetic diversity. Disease-related SNPs are sometimes also called disease-related loci.
[0082] "Polymorphic alleles" and "polymorphic loci" refer to alleles and loci, respectively, that have different genotypes among individuals of a given species. Some examples of polymorphic alleles include single nucleotide polymorphisms, short tandem repeats, deletions, duplications, and inversions.
[0083] A "polymorphic region" refers to a specific nucleotide found in a polymorphic region that differs between individuals.
[0084] "Mutation" refers to a change in the native nucleic acid sequence or reference nucleic acid sequence, and includes insertions, deletions, duplications, translocations, substitutions, reading frame shift mutations, silent mutations, nonsense mutations, missense mutations, point mutations, base transposition mutations, base conversion mutations, reversion mutations, microsatellite changes, etc. In some embodiments, the amino acid sequence encoded by the nucleic acid sequence has at least one amino acid change from the native sequence.
[0085] An "allele" refers to a group of genes that occupy a specific locus.
[0086] Both "genetic data" and "genotype data" refer to descriptions of aspects of the genome of one or more individuals. The term may refer to a locus or set of loci, a portion or entire sequence, a portion or entire chromosome, or the entire genome. The term may refer to the identity of one or more nucleotides. It may also refer to a nucleotide sequence, nucleotides at different locations in the genome, or combinations thereof. However, genotype data is typically computer-based, and the physical nucleotides of a given sequence can also be considered as chemically encoded genetic data. Genotype data may be described as genotype data "on," "of," "in," "from," or "about" an individual. Genotype data may also refer to the measurement output obtained from a genotyping platform that measures genetic material.
[0087] "Genetic material" and "genetic sample" refer to substances (such as tissue or blood) derived from one or more individuals that contain DNA or RNA.
[0088] "Confidence level" refers to the statistical likelihood that a called SNP, allele, set of alleles, determined copy number of a chromosome or chromosomal segment, or diagnosis of disease presence or absence, correctly represents the actual genetic state of the individual.
[0089] "Ploidy calling" and "chromosome copy number calling" or "copy number calling" (CNC) may refer to the process of determining the quantity and / or chromosomal identity of one or more chromosomes or chromosome segments present in a given cell.
[0090] "Aneuploidy" refers to a condition in which a cell has an error in the number of chromosomes (for example, an error in the number of complete chromosomes, or an error in the number of chromosome segments (such as the presence of deletions or duplications of chromosome segments)). In the case of human somatic cells, this term may refer to a cell that does not have 22 pairs of autosomes and 1 pair of sex chromosomes. In the case of human gametes, this term may refer to a cell that does not have one of the 23 chromosomes. In the case of monochromosome type, this term may refer to the presence of three or more or one or fewer homologous but non-identical chromosome copies, or the presence of two chromosome copies from the same parent. In some embodiments, deletions of chromosome segments are microdeletions.
[0091] "Ploidy" refers to the quantity and / or chromosomal identity of one or more chromosomes or chromosomal segments in a cell.
[0092] The term "chromosome" can refer to a single copy of a chromosome. This term means a single molecule of DNA, of which there are 46 in a normal somatic cell (e.g., "maternally inherited chromosome 18"). A chromosome can also refer to a specific chromosome type, of which there are 23 in a normal human somatic cell (e.g., "chromosome 18").
[0093] "Chromosomal identity" can refer to the chromosome reference number, or chromosome type. A normal human being has 22 numbered autosomes and two sex chromosomes. This term can also refer to the parental origin of a chromosome. It can also refer to a specific chromosome inherited from a parent. It can also refer to other features used for chromosome identification.
[0094] "Allele data" refers to a set of genotype data relating to one or more alleles. This term may also refer to phase-determined haplotype data. It may also refer to SNP identity. Furthermore, it may refer to DNA sequence data, including insertions, deletions, repeats, and mutations. This term may include the parental origin of each allele.
[0095] "Allele state" refers to the actual state of a gene in a set of one or more alleles. This term may also refer to the actual state of a gene as described by allele data.
[0096] The "allele number" refers to the number of sequences mapped to a particular locus, or, if the locus is polymorphic, the number of sequences mapped to each allele. When each allele is measured in binary, the allele number is an integer. When alleles are measured probabilistically, the allele number may be a fraction.
[0097] The "allele number probability" refers to the number of sequences that are likely to map to a set of alleles at a particular locus or polymorphic locus, combined with the probability of mapping. If the mapping probability of each measured sequence is binary (0 or 1), the allele number corresponds to the allele number probability. In some embodiments, the allele number probability may be binary. In some embodiments, the allele number probability may be set to be equal to the DNA measurement.
[0098] The "allele distribution" or "allele number distribution" refers to the relative quantity of each allele present in each locus within a set of loci. The allele distribution can refer to an individual, a sample, or a set of measurements performed on a sample. In the context of digital allele measurements such as sequencing, this allele distribution refers to the number or probability of reads mapping to a particular allele for each allele in a set of polymorphic loci. In the context of analog allele measurements such as SNP arrays, this allele distribution refers to allele intensity and / or allele ratio. Allele measurements may be treated as probabilities, i.e., as a fraction between 0 and 1 indicating the likelihood that a given allele exists in a given sequence read, or in binary, i.e., in which any given read is considered either 0 or 1 as a copy of a particular allele.
[0099] An "allele distribution pattern" refers to a set of allele distributions that differ depending on the context (e.g., the parental context). A certain allele distribution pattern may refer to a certain ploidy state.
[0100] "Allele bias" refers to the degree to which the measured ratio of alleles present at heterozygous loci differs from the ratio in the original DNA or RNA sample. The degree of allele bias at a particular locus is equal to the ratio of the observed allele ratio present at that locus to the ratio of alleles present at that locus in the original DNA or RNA sample. Allele bias can be due to amplification bias, purification bias, or several other phenomena that affect different alleles in various ways.
[0101] In the case of SNVs, "allelelic disequilibrium" refers to the proportion of abnormal DNA. This proportion is typically measured using the diversity allele frequency (number of diversity alleles present at a given locus / total number of alleles present at that locus). Since the difference in the amount of two homologs in a tumor is similar, the proportion of abnormal DNA in CNVs is measured using the mean allelelic disequilibrium (AAI). The mean allelelic disequilibrium is defined as |(H1-H2)| / (H1+H2), where Hi is the average copy number of homolog i in the sample, and Hi / (H1+H2) is the fractional amount of homolog i, i.e., the homolog ratio. The maximum homolog ratio is the homolog ratio of the most abundant homolog.
[0102] The "trial dropout rate" is the percentage of SNPs without reads, estimated using all SNPs.
[0103] The "single allele dropout (ADO) rate" is the percentage of SNPs that contain only one allele, estimated using only heterozygous SNPs.
[0104] "Primers" and "PCR probes" refer to single nucleic acid molecules (DNA molecules, DNA oligomers, etc.) or assemblies of identical or nearly identical nucleic acid molecules (DNA molecules, DNA oligomers, etc.). A primer contains a region designed to hybridize to a target locus (e.g., a target polymorphic or non-polymorphic locus) or a general-purpose priming sequence. A primer may also contain a priming sequence designed to enable PCR amplification. A primer may also contain a molecular barcode. A primer may contain different random regions for each individual molecule.
[0105] A "primer library" refers to a collection of two or more primers. In various embodiments, this library contains at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different primers. In various embodiments, the library includes at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different primer pairs, each primer pair including a forward test primer and a reverse test primer, and each test primer pair is hybridized to a target locus. In some embodiments, the primer library includes at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different individual primers. Each individual primer is hybridized to a different target locus, and these individual primers are not part of a primer pair. In some embodiments, the library includes both (i) primer pairs and (ii) individual primers that are not part of a primer pair (e.g., general-purpose primers).
[0106] "Different primers" refers to primers that are not identical.
[0107] "Different pools" refers to pools that are not identical.
[0108] "Different target seating positions" refers to target seating positions that are not identical.
[0109] "Different amplification products" refers to amplification products that are not identical.
[0110] A "hybrid capture probe" refers to any modifiable nucleic acid sequence created by various methods such as PCR or direct synthesis, intended to be complementary to a single strand of a specific target DNA sequence in the sample. An exogenous hybrid capture probe may be added to the prepared sample and hybridized by a denaturation re-annealing method to form a double-stranded exogenous-endogenous fragment. These double-stranded fragments may then be physically separated from the sample by various means.
[0111] A "sequence read" refers to data representing the base sequence of nucleotides measured, for example, using a clonal sequencing method. Clonal sequencing may produce sequence data representing a single original DNA molecule, its clones, or its clusters. A sequence read may also have an associated quality score at each base position in the sequence, indicating the probability that the nucleotide was correctly called.
[0112] Sequence read mapping is a method for determining the original location of a sequence read within the genome sequence of a specific organ. The original location of a sequence read is based on the similarity of its nucleotide sequence and genome sequence.
[0113] "Constant copy error" and "constant aneuploidy (MCA)" refer to aneuploidy in which a single cell contains two identical or nearly identical chromosomes. This type of aneuploidy can occur during meiosis when gamete formation takes place, and is sometimes called meiotic nondisjunction error. This type of error can also occur during mitosis. Constant triploidy sometimes refers to a situation where three copies of a chromosome exist in a single individual, two of which are identical.
[0114] "Partially identical copy errors" and "non-identical aneuploidy (UCA)" refer to aneuploidy conditions in which a single cell contains two chromosomes from the same parent that are homologous but may not be identical. This type of aneuploidy can occur during meiosis and is sometimes called a meiotic error. Partially identical trisomal chromosomes can refer to a case where an individual has three copies of a chromosome, two of which are from the same parent and are homologous but not identical. Partially identical trisomal chromosomes can also refer to a case where two homologous chromosomes from one parent exist, and some parts of their chromosome segments are identical, while other segments are simply homologous.
[0115] "Homologous chromosomes" refer to chromosomal copies that contain the same set of genes that pair up normally during meiosis.
[0116] "Identical chromosomes" refer to chromosomal copies that contain the same set of genes, and each gene has the same set of alleles that are identical or nearly identical.
[0117] Allele dropout (ADO) refers to a situation where, in a set of base pairs on homologous chromosomes of a particular allele, at least one of the base pairs is not detected.
[0118] Locus loss (LDO) refers to a situation where, within a set of base pairs on homologous chromosomes of a particular allele, neither base pair is detected.
[0119] "Homozygous" refers to having similar alleles at corresponding chromosomal loci.
[0120] "Heterozygous" refers to having alleles that are not similar to each other in their corresponding chromosomal loci.
[0121] "Heterozygosity rate" refers to the percentage of individuals in a population who have a heterozygous allele at a particular locus. Heterozygosity rate may also refer to the expected or measured ratio of alleles present at a particular locus in an individual or DNA or RNA sample.
[0122] A "chromosomal region" refers to a segment of a chromosome, or an entire chromosome.
[0123] "Chromosome segmentation" refers to a portion of a chromosome, and can range in size from a single base pair to a complete chromosome.
[0124] The term "chromosome" refers to a complete chromosome, or a segment or part of a chromosome.
[0125] "Copy" refers to the copy number of a chromosome segment. This term can refer to identical copies of a chromosome segment, or copies that are not identical but homologous. In non-identical but homologous copies, each copy of the chromosome segment contains substantially similar locus sets, with one or more alleles being different. In some cases of aneuploidy, such as M2 copy errors, a chromosome segment may have some copies that are identical and some that are not.
[0126] A "haplotype" typically refers to a combination of alleles present at multiple loci on the same chromosome. A haplotype can range from at least two loci to a complete chromosome, depending on the number of recombination events that have occurred between a given set of loci. A haplotype can also refer to a statistically related set of SNPs on a single chromatid.
[0127] "Haplotype data," "phase-determined data," and "ordered gene data" refer to data derived from a single chromosome or chromosome segment in a diploid or polyploid organism, such as isolated maternal or paternal chromosome copies in a diploid genome.
[0128] "Phase determination" refers to the process of determining an individual's haplotype gene data in a given, unordered, diploid (or polyploidy) gene dataset. The term may also refer to the process of determining, with respect to a set of alleles found on a single chromosome, which of two genes in a given allele is associated with which of the individual's two homologous chromosomes.
[0129] "Phase-determined data" refers to genetic data in which one or more haplotypes have been determined.
[0130] A "hypothesis" refers to a possible state (for example, the degree to which copy number overpopulation is likely to occur in a first homologous chromosome or chromosome segment compared to a second homologous chromosome or chromosome segment, possible deletions, possible duplications, possible ploidy states in a set of one or more chromosomes or chromosome segments, possible allele states in a set of one or more loci, possible paternity relationships, possible amounts of DNA, RNA, fetal fractions, and genetic material derived from a set of loci in a set of one or more chromosomes or chromosome segments). A genetic state may be associated with the relative likelihood that each element of the hypothesis is true in relation to the other elements of the hypothesis, or the relative likelihood that the hypothesis as a whole is true. This set of possibilities may contain one or more elements.
[0131] The "copy number hypothesis" and the "ploidy status hypothesis" refer to hypotheses about the copy number of a chromosome or chromosome segment in an individual. These terms may also refer to hypotheses about the identity of each chromosome (including the original parents of each chromosome and which of the two chromosomes of those parents is present in the individual). These terms may also refer to hypotheses about the genetic correspondence between a chromosome or chromosome segment originating from a related individual, if such individuals exist, and a particular chromosome in that individual.
[0132] A “related individual” refers to any individual that is genetically related to the individual in question and shares a haplotype block with it. In some contexts, a related individual may be the genetic parent of the individual in question, or any genetic material derived from a parent (sperm, polar body, embryo, fetus, or offspring). This term may also refer to siblings, parents, or grandparents.
[0133] The term "sibling" refers to any individual whose genetic parents are the same as the individual in question. In some aspects, the term may refer to a newborn child, embryo, fetus, or one or more cells derived from a newborn child, embryo, or fetus. Siblings may also refer to a haploid individual derived from either parent (e.g., sperm, polar body, or other collection of haplotype genetic material). An individual may be considered its own sibling.
[0134] The term "child" may refer to an embryo, blastomeres, or fetus. In the embodiments disclosed herein, the concepts described apply equally to an individual that is a newborn child, fetus, embryo, or collection of cells derived therefrom. When the term "child" is used, it may simply mean that the individual referred to as a child is a genetic offspring of its parents.
[0135] "Fetal" refers to "of that fetus" or "a placental region genetically similar to that fetus." Certain parts of a pregnant woman's placenta are genetically similar to the fetus, and free-floating fetal DNA found in the mother's blood may originate from a portion of the placenta that has a genotype matching that of the fetus. Half of the genetic information of a fetus's chromosomes is inherited from the fetus's mother. In some aspects, DNA obtained from fetal cells that originates from chromosomes inherited from the mother is considered "fetal" rather than "maternal."
[0136] "Fetal DNA" refers to DNA that was originally part of a cell whose genotype was essentially the same as that of a fetus.
[0137] "Maternal DNA" refers to DNA that was originally part of a cell whose genotype was essentially the same as that of the mother.
[0138] "Parents" refers to the genetic mother or father of an individual. An individual typically has two parents, a mother and a father, but this is not always the case in cases of genetic chimerism or chromosomal chimerism. Parents may sometimes be considered individuals.
[0139] "Parental context" refers to the genetic state of a particular SNP in each of the two related chromosomes of one or both of the two parents in question.
[0140] "Maternal plasma" refers to the plasma portion of blood derived from a pregnant woman.
[0141] A "clinical judgment" refers to any decision on whether or not to take an action that has consequences affecting the health or survival of an individual. A clinical judgment may also refer to taking further tests, having an abortion, maintaining a pregnancy, taking action to mitigate an undesirable phenotype, or taking action to prepare for a particular phenotype.
[0142] "Diagnostic box" means one machine or combination of such machines designed to perform one or more aspects of the methods disclosed herein. In one embodiment, the diagnostic box may be provided in terms of patient treatment. In one embodiment, the diagnostic box may perform amplification of the desired effect and subsequent sequencing. In one embodiment, the diagnostic box may function independently or with the assistance of a technician.
[0143] "Informatics-based methods" refer to methods that rely heavily on statistics to interpret large amounts of data. In the context of prenatal diagnosis, this term refers to methods designed to determine the ploidy status of one or more chromosomes or chromosomal segments, the allele status of one or more alleles, or paternity by statistically inferring the most likely status, rather than directly measuring the state physically, when large amounts of genetic data (e.g., from molecular arrays or sequencing) are available. In one aspect of this disclosure, the informatics-based technology may be what is disclosed in this patent application. In one aspect of this disclosure, the informatics-based technology may be PARENTAL SUPPORT (trademark: a chromosomal aneuploidy screening technology developed by Gen Security Network, Inc.).
[0144] "Primary gene data" refers to the analog intensity signals output by a genotyping platform. In the context of SNP arrays, primary gene data refers to the intensity signals before any genotype calling is performed. In the context of sequencing, primary gene data refers to chromatogram-like analog measurements output by a sequencing instrument before the identity of any base pair is determined and before the sequence is mapped to the genome.
[0145] "Secondary gene data" refers to processed gene data output by a genotyping platform. In the context of SNP arrays, secondary gene data refers to allele retrieval performed by software associated with the SNP array reader, which retrieves whether a particular allele is present in the sample. In the context of sequencing, secondary gene data refers to the base pair identity of the determined sequence. It can also refer to the base pair identity when the sequence is mapped to the genome.
[0146] "Preferential enrichment of DNA corresponding to a locus or preferential enrichment of DNA present at a locus" refers to any method that makes the percentage of DNA molecules corresponding to that locus in the enriched DNA mixture higher than the percentage of DNA molecules corresponding to that locus in the unenriched DNA mixture. This method may include selective amplification of DNA molecules corresponding to a locus. This method may include removal of DNA molecules that do not correspond to that locus. This method may include a combination of several methods. The degree of enrichment is defined as the percentage of DNA molecules corresponding to that locus in the enriched mixture divided by the percentage of DNA molecules corresponding to that locus in the unenriched mixture. Preferential enrichment may be performed for multiple loci. In some aspects of this disclosure, the degree of enrichment is greater than 20, 200, or 2,000. When preferential enrichment is performed for multiple loci, the degree of enrichment may refer to the average degree of enrichment of all loci in the set of loci.
[0147] "Amplification" refers to a method of increasing the copy number of DNA or RNA molecules.
[0148] "Selective amplification" may refer to a method of increasing the copy number of a particular DNA (or RNA) molecule, or the copy number of a DNA (or RNA) molecule corresponding to a particular region of DNA (or RNA). The term may also refer to a method of increasing the copy number of a particular target DNA (or RNA) molecule or region more than the increase of non-target DNA (or RNA) molecules or regions. Selective amplification may also be a method of preferential enrichment.
[0149] A "generic priming sequence" refers to a DNA (or RNA) sequence that may be added to a population of target DNA (or RNA) molecules, for example, by ligation, PCR, or ligation-mediated PCR. Once a generic priming sequence is added to a population of target molecules, this population can be amplified using a pair of amplification primers with primers specific to this generic priming sequence. Generic priming sequences are usually unrelated to the target sequence.
[0150] A "general-purpose ligator," "ligation ligator," or "library label" is a nucleic acid molecule containing a general-purpose priming sequence that can be covalently attached to the 5' and 3' ends of a population of target double-stranded nucleic acid molecules. The addition of the ligator supplies the general-purpose priming sequence to the 5' and 3' ends of a target population that can initiate PCR amplification, allowing all molecules from this population to be amplified using a pair of amplification primers.
[0151] "Objectification" refers to a method used to selectively amplify or preferentially enrich DNA (or RNA) molecules that correspond to a set of loci in a DNA (or RNA) mixture.
[0152] A "joint distribution model" refers to a model that defines the probability of an event defined by several random variables, assuming that the probabilities of the variables are combined in the same probability space. In some embodiments, a degenerate case where the probabilities of the variables are not combined may be used.
[0153] "Cancer-related genes" refer to genes associated with changes in cancer risk or prognosis. Examples of cancer-related genes that promote cancer include oncogenes, genes that increase cell proliferation, invasion, or metastasis, genes that inhibit apoptosis, and genes that promote angiogenesis. Examples of cancer-related genes that inhibit cancer include, but are not limited to, tumor suppressor genes, genes that inhibit cell proliferation, invasion, or metastasis, genes that promote apoptosis, and genes that inhibit angiogenesis.
[0154] "Estrogen-related cancers" refer to cancers regulated by estrogen. Examples of estrogen-related cancers include, but are not limited to, breast cancer and ovarian cancer. Her2 is overexpressed in many estrogen-related cancers (U.S. Patent No. 6,165,464, which is incorporated herein by reference in its entirety).
[0155] "Androgen-related cancers" refer to cancers that are regulated by androgens. An example of an androgen-related cancer is prostate cancer.
[0156] "Higher than normal expression levels" means that the expression level of mRNA or protein is higher than the average expression level of a control subject (for example, a subject without cancer or other disease). In various aspects, this higher than normal expression level is at least 20, 40, 50, 75, 90, 100, 200, 500, or 1000% higher than the expression level in the control subject.
[0157] "Lower than normal expression levels" means that the expression level of mRNA or protein is lower than the average expression level of a control subject (e.g., a subject without cancer or other disease). In various embodiments, this higher-than-normal expression level is at least 20, 40, 50, 75, 90, 95, or 100% lower than the expression level in the control subject. In some embodiments, mRNA or protein expression is undetectable.
[0158] "Modifying expression or activity" refers, for example, to increasing or decreasing the expression or activity of a protein or nucleic acid sequence compared to a control condition. In some embodiments, the modulation of expression or activity is an increase or decrease of at least 10, 20, 40, 50, 75, 90, 100, 200, 500, or 1000%. In various embodiments, transcription, translation, mRNA or protein stability, or the binding of mRNA or protein to other molecules in the body are modulated by the treatment. In some embodiments, the amount of mRNA is determined by standard Northern blot analysis, and the amount of protein is determined by standard Western blot analysis (e.g., the analysis described herein or the analysis described, for example, in Ausubel et al. (Current Protocols in Molecular Biology, John Wiley & Sons, New York, July 11, 2013, which is incorporated herein in its entirety by reference)). In one embodiment, the amount of protein is determined by measuring the amount of enzyme activity using standard methods. In other preferred embodiments, the amount of mRNA, protein, or enzyme activity is greater than that of control cells not expressing the functional form of the protein (e.g., cells that are homozygous in the case of nonsense mutations), but 20 times, 10 times, 5 times, or 2 times or less. In yet other embodiments, the amount of mRNA, protein, or enzyme activity is greater than that of control cells (e.g., non-cancer cells, cells not exposed to conditions that induce abnormal cell proliferation or inhibit apoptosis, cells derived from subjects without the disease or disorder of the interest), but 20 times, 10 times, 5 times, or 2 times or less.
[0159] "A dose sufficient to modulate mRNA or protein expression or activity" refers to the amount of treatment administered to a subject that increases or decreases mRNA or protein expression or activity. In some embodiments, for compounds that reduce expression or activity, the modulation is a reduction of at least 10%, 30%, 40%, 50%, 75%, or 90% in the treated subject compared to the same subject before administration of the inhibitor or to an untreated control subject. In some embodiments, for compounds that increase expression or activity, the mRNA or protein expression or activity level in the treated subject is 1.5, 2, 3, 5, 10, or 20 times higher than the expression or activity level in the same subject before administration of the modulator or to an untreated control subject.
[0160] In some embodiments, compounds may directly or indirectly regulate the expression or activity of RNA or proteins. For example, a compound may indirectly regulate the expression or activity of a target mRNA or protein by regulating the expression or activity of a molecule (e.g., nucleic acid, protein, signaling molecule, growth factor, cytokine, or chemokine) that directly or indirectly affects the expression or activity of the target mRNA or protein. In some embodiments, the compounds may inhibit cell division or induce apoptosis. Examples of these compounds in the therapeutic agent include unpurified or purified proteins, antibodies, synthetic organic molecules, natural organic molecules, nucleic acid molecules, and components thereof. In combination therapy, these compounds may be administered simultaneously or sequentially. Examples of these compounds include signaling inhibitors.
[0161] "Purification" refers to separation from other naturally occurring components. Generally, a factor is substantially pure if it does not contain at least 50% by weight of proteins, antibodies, and naturally occurring organic molecules associated with it in its natural state. In some embodiments, the factor is at least 75% by weight, 90% by weight, or 99% by weight pure. Substantially pure factors may be obtained by chemosynthesis, isolation of the factor from its natural source, or production of the factor in recombinant host cells that do not naturally produce it. Proteins and small molecules may be purified by those skilled in the art using standard techniques, for example, those described in Ausubel et al. (Current Protocols in Molecular Biology, John Wiley & Sons, New York, July 11, 2013, which is incorporated herein by reference in its entirety). In some embodiments, the factor is at least 2-fold, 5-fold, or 10-fold pure of the starting material in measurements using polyacrylamide gel electrophoresis, column chromatography, optical density, HPLC analysis, or Western chemistry (Ausubel et al., supra). Purification methods include, for example, immunoprecipitation, column chromatography (e.g., immunoaffinity chromatography), magnetic bead immunoaffinity purification, and panning with plate-conjugated antibodies.
[0162] Other features and advantages of the present invention will become apparent from the following detailed description and claims.
[0163] This patent, i.e., the application documents, include at least one color drawing. A copy of this patent or patent application publication with the color drawing will be provided by the Japan Patent Office upon request and payment of the necessary fees.
[0164] The embodiments disclosed herein will be further described with reference to the accompanying drawings. In some drawings, similar structures are denoted by the same reference numerals. The drawings shown herein are not necessarily to scale and are rather intended to emphasize the principles of the embodiments disclosed herein. [Brief explanation of the drawing]
[0165] [Figure 1A-1D] This graph shows the distribution for different copy number hypotheses, calculated by dividing the test statistic S by T (number of SNPs) ("S / T") as the number of SNPs increases, with a read depth (DOR) of 500 and a tumor fraction of 1%.
[0166] [Figure 2A-2D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 500 and the tumor fraction is 2%, in response to an increase in the number of SNPs.
[0167] [Figure 3A-3D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 500 and the tumor fraction is 3%, in response to an increase in the number of SNPs.
[0168] [Figure 4A-4D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 500 and the tumor fraction is 4%, in response to an increase in the number of SNPs.
[0169] [Figures 5A-5D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 500 and the tumor fraction is 5%, in response to an increase in the number of SNPs.
[0170] [Figures 6A-6D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 500 and the tumor fraction is 6%, in response to an increase in the number of SNPs.
[0171] [Figures 7A-7D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 1000 and the tumor fraction is 0.5%, in response to an increase in the number of SNPs.
[0172] [Figures 8A-8D]This graph shows the S / T distribution for different copy number hypotheses when the DOR is 1000 and the tumor fraction is 1%, as the number of SNPs increases.
[0173] [Figures 9A-9D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 1000 and the tumor fraction is 2%, as the number of SNPs increases.
[0174] [Figure 10A-10D] This graph shows the S / T distribution for different copy number hypotheses as the number of SNPs increases, with a DOR of 1000 and a tumor fraction of 3%.
[0175] [Figure 11A-11D] This graph shows the S / T distribution for different copy number hypotheses as the number of SNPs increases, with a DOR of 1000 and a tumor fraction of 4%.
[0176] [Figures 12A-12D] This graph shows the S / T distribution for different copy number hypotheses as the number of SNPs increases, with a DOR of 3000 and a tumor fraction of 0.5%.
[0177] [Figures 13A-13D] This graph shows the S / T distribution for different copy number hypotheses when the DOR is 3000 and the tumor fraction is 1%, as the number of SNPs increases.
[0178] [Figure 14] This table shows the sensitivity and specificity for detecting six types of microdeletion syndromes.
[0179] [Figure 15A-15C]This graph displays euploidy. The x-axis represents the linear position of the polymorphic locus of an individual along the chromosome, and the y-axis represents the number of reads for allele A as a fraction of the total number of reads (A+B) for alleles. The maternal genotype and fetal genotype are shown to the right of the plot. Each plot is colored according to the maternal genotype. For example, red represents maternal genotype AA, blue represents maternal genotype BB, and green represents maternal genotype AB. Figure 15A is a plot when two chromosomes are present and the fetal cell-free DNA fraction is 0%. This plot represents a pattern derived from a non-pregnant woman, and therefore all of its genotypes are maternal. Consequently, the allele groups are concentrated at 1 (allele of AA), 0.5 (allele of AB), and 0 (allele of BB). Figure 15B is a plot when two chromosomes are present and the fetal fraction is 12%. The position of the allele points is shifted up or down along the y-axis by the fetal alleles included in the fraction of allele A reads. Figure 15C is a plot of the case where two chromosomes are present and the fetal fraction is 26%. This pattern includes two red peripheral bands and two blue peripheral bands, and three green bands are easily visible.
[0180] [Figures 16A-16B] This graph shows the 22q11.2 deletion syndrome. Figure 16A shows the case of a mother who carries the 22q11.2 deletion (indicated by the absence of SNPs in green A and B). Figure 16B shows the case of a fetus who inherited the 22q11 deletion from the parent (indicated by one red peripheral band and one blue peripheral band). The x-axis represents the linear position of the SNP, and the y-axis represents the fractionation of allele A reads in the total number of reads. Each point represents a single SNP locus.
[0181] [Figure 17] This is a graph of the cat cry deletion syndrome inherited from the mother (indicated by the presence of two central green bands instead of three green bands). The x-axis represents the linear position of the SNP, and the y-axis represents the fractionation of allele A reads in the total number of reads. Each point represents a single SNP locus.
[0182] [Figure 18] This is a graph of Wolff-Hirschhorn deletion syndrome (indicated by one red and one blue peripheral band). The x-axis represents the linear position of the SNP, and the y-axis represents the fractionation of allele A reads in the total number of reads. Each point represents a single SNP locus.
[0183] [Figure 19A] This is a graphical representation of the X chromosome spike-in experiment, showing the excess copies of a chromosome or chromosome segment. These plots show different amounts of paternal DNA mixed with daughter DNA, with paternal DNA at 16% (Figure 19A), 10% (Figure 19B), 1% (Figure 19C), and 0.1% (Figure 19D), respectively. The x-axis represents the linear position of SNPs on the X chromosome, and the y-axis represents the fraction of reads for allele M in the total number of reads (M+R). Each point indicates a single SNP locus containing allele M or R. [Figure 19B] Same as above. [Figure 19C] Same as above. [Figure 19D] Same as above.
[0184] [Figures 20A-20B] Figure 20A shows a graph of the false negative rate using haplotype data, and Figure 20B shows a graph of the false negative rate without using haplotype data.
[0185] [Figures 21A-21B] Figure 21A shows the false positive rate at p=1% using haplotype data, and Figure 21B shows the false positive rate at p=1% without using haplotype data.
[0186] [Figures 22A-22B] Figure 22A shows the false positive rate at p=1.5% using haplotype data, and Figure 22B shows the false positive rate at p=1.5% without using haplotype data.
[0187] [Figures 23A-23B] Figure 23A shows the false positive rate at p=2% using haplotype data, and Figure 23B shows the false positive rate at p=2% without using haplotype data.
[0188] [Figures 24A-24B] Figure 24A shows the false positive rate at p=2.5% using haplotype data, and Figure 24B shows the false positive rate at p=2.5% without using haplotype data.
[0189] [Figures 25A-25B] Figure 25A shows the false positive rate at p=3% using haplotype data, and Figure 25B shows the false positive rate at p=3% without using haplotype data.
[0190] [Figure 26] This is the false positive rate table for the first simulation.
[0191] [Figure 27] This is the false negative rate table for the first simulation.
[0192] [Figures 28A-28C] Figure 28A is a graph of the reference count (number of alleles (e.g., allele "A")) obtained by dividing the normal (non-cancerous) cell line by the total number of loci. Figure 28B is a graph of the reference count obtained by dividing the deletion by the total number for cancer cell lines. Figure 28C is a graph of the reference count divided by the total count for a mixture of DNA derived from the above normal and cancer cell lines.
[0193] [Figure 29]This graph shows the ratio of baseline counts to total counts for plasma samples from stage IIa breast cancer patients, where the tumor fraction was estimated at 4.33% (4.33% of the DNA was tumor cells). The green portion of the graph represents the region where no CNVs are present. The blue and red portions of the graph represent the region where CNVs are present and the measured allele ratio is clearly separated from the expected allele ratio of 0.5. Blue indicates one haplotype, and red indicates another haplotype. Approximately 636 heterozygous SNPs were analyzed in the CNV region.
[0194] [Figure 30] This graph shows the ratio of baseline counts to total counts for plasma samples from stage IIb breast cancer patients with an estimated tumor fraction of 0.58%. The green areas in this graph represent regions where CNVs are absent. The blue and red areas in this graph indicate regions where CNVs are present, but the measured allele ratio is not clearly separated from the expected allele ratio of 0.5. In this analysis, 86 heterozygous SNPs were analyzed in the CNV regions.
[0195] [Figure 31A-31B] This graph shows the maximum likelihood estimate of the tumor fraction. The maximum likelihood estimate is indicated by the peak of the graph, which is 4.33% in Figure 31A and 0.58% in Figure 31B.
[0196] [Figures 32A-32B] Figure 32A is a graph comparing the log odds ratios of various possible tumor fractions for high tumor fraction samples (4.33%) and low tumor fraction samples (0.58%). When the log odds ratio is less than 0, the polyploid hypothesis is the most likely. When the log odds ratio is greater than 0, the presence of CNV is the most likely. Figure 32B is a graph for low tumor fraction samples (0.58%), showing the probability of deletion divided by the probability of no deletion for various possible tumor fractions.
[0197] [Figure 33]This is a graph showing the log odds ratios of various possible tumor fractions for a low tumor fraction sample (0.58%). Figure 33 is an enlarged graph of the low tumor fraction sample in Figure 32A.
[0198] [Figure 34] This is a graph showing the detection limits of single nucleotide polymorphisms in tumor biopsy materials using three different methods described in Example 6.
[0199] [Figure 35] This is a graph showing the detection limits of single nucleotide polymorphisms in plasma samples using three different methods described in Example 6.
[0200] [Figures 36A-36B] This is a graph analyzing genomic DNA (Figure 36A) or DNA derived from a single cell (Figure 36B) using a library of approximately 28,000 primers designed to detect copy number variations (CNVs). The presence of CNV is indicated by two bands instead of one in the center. The x-axis represents the linear position of SNPs, and the y-axis represents the fraction of reads of allele A in the total number of reads.
[0201] [Figures 37A-37B] This is a graph analyzing genomic DNA (Figure 37A) or DNA derived from a single cell (Figure 37B) using a library of approximately 3,000 primers designed to detect copy number variations (CNVs). The presence of CNV is indicated by two bands instead of one in the center. The x-axis represents the linear position of SNPs, and the y-axis represents the fraction of reads of allele A in the total number of reads.
[0202] [Figure 38] This is a graph showing the uniformity in the DOR of approximately 3,000 loci.
[0203] [Figure 39] This is a table comparing the evaluation criteria for false calls for genomic DNA and DNA derived from a single cell.
[0204] [Figure 40] This graph shows the error rates for translocation and conversion mutations.
[0205] [Figures 41a-41d] This graph shows the sensitivity of CoNVERGe determined by PlasmArt. (a) shows the correlation between the mean allele mismatch rate (AAI) calculated by CoNVERGe and the actual input fraction for PlasmArt samples containing DNA from 22q11.2 deletion and matching normal cell lines. (b) shows the correlation between the calculated AAI and the actual tumor DNA input for PlasmArt samples containing DNA from HCC2218 breast cancer cells, including CNVs on chromosomes 2p and 2q and matching normal HCC2218BL cells, with a tumor DNA fraction of 0-9.09%. (c) shows the correlation between the calculated AAI and the actual tumor DNA input for PlasmArt samples containing DNA from HCC1954 breast cancer cells, including CNVs on chromosomes 1p and 1q and matching normal HCC1954BL cells, with a tumor DNA fraction of 0-5.66%. (d) is a plot of the allele frequencies of HCC1954 cells used in (c). In (a), (b), and (c), the data points and error bars represent the mean and standard deviation (SD) of 3 to 8 iterations, respectively.
[0206] [Figure 42] This section provides an example of details regarding the PlasmArt standard, and includes a graph at the bottom showing the dimensional distribution of the fragments.
[0207] [Figure 43A] The left panel shows the dilution curves of results obtained from PlasmArt synthetic ctDNA standard samples for validating the microdeletion and cancer panels. The right panel shows the maximum likelihood estimates of tumors for the DNA fraction as a plot of odds ratios. [Figure 43B-C] Figure 43B is a plot in which a transition event was detected. Figure 43C is a plot in which a change event was detected.
[0208] [Figure 44] Plots showing CNV of different chromosomal regions for various samples with different percentages of ctDNA.
[0209] [Figure 45] Plots showing CNV of different chromosomal regions for various ovarian cancer samples with different amounts of ctDNA (%).
[0210] [Figure 46] Table showing the percentage of breast or lung cancer patients with SNV in ctDNA, or combinations of SNV and / or CNV.
[0211] [Figure 47] Graph showing tumor-specific SNV and / or CNV (%) contained in plasma samples in breast cancer at different stages, and the associated data table on its right side.
[0212] [Figure 48] Graph showing tumor-specific SNV and / or CNV (%) contained in plasma samples in different sub-stages of breast cancer, and the associated data table on its right side.
[0213] [Figure 49] Graph showing tumor-specific SNV and / or CNV (%) contained in plasma samples in lung cancer at different stages, and the associated data table on its right side.
[0214] [Figure 50] Graph showing tumor-specific SNV and / or CNV (%) contained in plasma samples in different sub-stages of lung cancer, and the associated data table on its right side.
[0215] [Figure 51A] Showing the histological findings / history of primary lung tumors analyzed for tumor clone and sub-clone heterogeneity. [Figure 51B]This table shows the identity of diversity allele frequencies (VAFs) of biopsied lung tumors obtained by whole-genome sequencing and AmpliSEQ assays.
[0216] [Figure 52] Using plasma-derived ctDNA, we demonstrate the identification of both clones and quasi-clones of single nucleotide addition (SNA) mutations to overcome tumor heterogeneity.
[0217] [Figure 53] This table compares VAF calling by AmpliSeq and mmPCR-NGS for detecting primary tumor SNVs, which are SNV mutations missed by AmpliSeq but identified in plasma-derived ctDNA.
[0218] [Figure 54] Figure 54A is a plot showing VAF (%) in primary lung tumors. Figure 54B is a linear regression plot comparing VAF calculated by AmpliSeq and VAF calculated by Natera.
[0219] [Figure 55] This graph shows pool 1 / 4 of the 84-plex PCR primer reactions for SNVs under limited primer concentrations.
[0220] [Figure 56] This graph shows pool 2 / 4 of the 84-plex PCR primer reactions for SNVs when primer concentrations are limited.
[0221] [Figure 57] This is a pool 3 / 4 graph of 84-plex PCR primer reactions for SNVs when primer concentrations are limited.
[0222] [Figure 58] This graph shows pool 4 / 4 of 84plex PCR primer reactions for SNVs when primer concentrations are limited.
[0223] [Figure 59] This plot shows the limit of detection (LOD) and read depth (DOR) when detecting SNV translocations and conversion mutations in an 84-plex PCR reaction over 15 PCR cycles.
[0224] [Figure 60] This plot shows the limit of detection (LOD) and read depth (DOR) when detecting SNV translocations and conversion mutations in an 84-plex PCR reaction with 20 PCR cycles.
[0225] [Figure 61] This plot shows the limit of detection (LOD) and read depth (DOR) when detecting SNV translocations and conversion mutations in an 84-plex PCR reaction over 25 PCR cycles.
[0226] [Figure 62] This plot compares the sensitivity of tumor cell genomic DNA and single-cell genomic DNA. The upper part shows tumor cell genomic DNA, and the lower part shows single-cell genomic DNA.
[0227] [Figure 63] Figure 63a shows the workflow for analyzing CNVs in various cancer samples using large-scale multi-stage PCR (mmPCR) assays targeting SNPs. Figures 63b to 63f show a comparison of CoNVERGe assays and microarray assays for breast cancer cell lines and matching normal cell lines.
[0228] [Figure 64] This shows a comparison of raw frozen (FF) breast cancer samples and formalin-fixed paraffin-embedded (FFPE) breast cancer samples with a matched control. Figures a-h compare CoNVERGe assays and microarray assays for breast cancer cell lines and a matched leptomeningeal gDNA control sample.
[0229] [Figure 65]This is a plot of allele frequencies detected in single cells, reflecting chromosomal copy number using the CoNVERGe assay. Figures 65a-65c show the analysis of three repeats in single breast cancer cells. Figure 65d shows the analysis of a B lymphocyte cell line without CNVs in the region of interest.
[0230] [Figure 66] This is a plot of allele frequencies detected in actual plasma samples using the CoNVERGe assay, reflecting chromosome copy number. Figure 66a shows a cell-free plasma DNA sample from stage II breast cancer and its corresponding tumor biopsy gDNA. Figure 66b shows a cell-free plasma DNA sample from advanced ovarian cancer and its corresponding tumor biopsy gDNA. Figure 66c is a table showing tumor heterogeneity determined by CNV detection in five advanced ovarian cancer plasma and corresponding tissue samples.
[0231] [Figure 67-1] This shows the changes in chromosomal location and mutations in breast cancer. [Figure 67-2] Same as above. [Figure 67-3] Same as above. [Figure 67-4] Same as above.
[0232] [Figure 68] The multiple allele frequencies (Figure 68A) and minority allele frequencies (Figure 68B) of SNPs used in the 3168 mM PCR reaction are shown.
[0233] [Figure 69] An example of a system design X00 useful for carrying out aspects of the present invention is shown.
[0234] [Figure 70]An example of a computer system for carrying out an aspect of the present invention is shown. While aspects are disclosed in these drawings, other aspects are also conceived, as discussed in the discussion. This disclosure provides typical and illustrative aspects, and is not limiting. Those skilled in the art can devise many other modifications and embodiments without departing from the principles and spirit of the aspects disclosed herein. [Modes for carrying out the invention]
[0235] In one aspect, the present invention generally relates to an improved method for determining the presence or absence of copy number diversity, such as deletions or duplications of chromosomal segments or complete chromosomes. The method is particularly useful for detecting small deletions or duplications. Such small deletions or duplications may be difficult to detect with high specificity and sensitivity using conventional methods due to the small amount of data available from the relevant chromosomal segments. The method includes an improved analytical method, an improved bioassay method, and a combination of the improved analytical method and the improved bioassay method. The method of the present invention can also be used to detect small percentage deletions or duplications present in cells or nucleic acid molecules being examined. This makes it possible to detect deletions or duplications before the onset of disease (e.g., precancerous stage) or in the early stages of disease (e.g., before deletions or duplications accumulate in a large number of diseased cells (e.g., cancer cells)). By more accurately detecting deletions or duplications associated with disease or disorder, improved methods can be provided for diagnosing, predicting, preventing, delaying, stabilizing, or treating that disease or disorder. Several types of deletions or duplications are known to be associated with cancer or certain mental or physical disorders.
[0236] In other aspects, the present invention generally relates to improved methods for detecting single nucleotide polymorphisms (SNVs), at least in particular. These improved methods include improved analytical methods, improved bioassay methods, and improved methods using a combination of improved analytical and improved bioassay methods. In certain exemplary embodiments, these methods are used to detect, diagnose, monitor, or stage cancer in samples where SNVs are present at very low concentrations (e.g., less than 10%, 5%, 4%, 3%, 2.5%, 2%, 1%, 0.5%, 0.25%, or 0.1%, such as in circulating free DNA samples). That is, in certain exemplary embodiments, these methods are particularly suitable for samples where relatively low percentage mutations or variants exist relative to the normal polymorphic alleles present at the gene locus. Finally, methods combining improved methods for detecting copy number diversity and improved methods for detecting single nucleotide polymorphisms are disclosed herein.
[0237] The success of treating diseases such as cancer often depends on early diagnosis, correct staging of the disease, selection of an effective treatment plan, and close monitoring to prevent or detect recurrence. In the case of cancer diagnosis, histological evaluation of tumor material obtained from tissue biopsy is often considered the most reliable method. However, due to the invasiveness of biopsy-based sampling, such sampling is difficult to perform in mass screenings and routine screenings. Therefore, the method of the present invention has the advantage of being non-invasive, relatively low-cost, and quick if desired. Target sequencing, which may be used with the method of the present invention, requires fewer reads than shotgun sequencing (e.g., millions of reads instead of 40 million), thereby reducing costs. Multiplex PCR and next-generation sequencing, which may be used, increase processing capacity and reduce costs.
[0238] In some embodiments, deletions, duplications, or single nucleotide variants are detected in an individual using the method described above. A sample derived from an individual containing cells or nucleic acids suspected of having deletions, duplications, or single nucleotide variants may be analyzed. In some embodiments, this sample is derived from tissue or organs (such as cells or tumors suspected of being cancerous) suspected of having deletions, duplications, or single nucleotide variants. Using the method of the present invention, deletions, duplications, or single nucleotide variants present in only one or a few cells in a mixture of cells having and not having deletions, duplications, or single nucleotide variants can be detected. In some embodiments, cell-free DNA or cell-free RNA derived from a blood sample from the individual is analyzed. In some embodiments, cell-free DNA or cell-free RNA is secreted by cells such as cancer cells. In some embodiments, cell-free DNA or cell-free RNA is released by necrotic or apoptotic cells such as cancer cells. Using the method of the present invention, deletions, duplications, or single nucleotide variants present in cell-free DNA or cell-free RNA at low percentages can be detected. In some embodiments, one or more cells derived from the embryo are examined.
[0239] In some embodiments, the methods are used in non-invasive or invasive prenatal testing of the fetus. These methods can be used to determine the presence or absence of deletions or duplications of chromosomal segments or entire chromosomes (e.g., deletions or duplications known to be associated with certain mental or physical disabilities, learning disabilities, or cancer). In some embodiments, non-invasive prenatal testing (NIPT) involves testing cells, cell-free DNA, or cell-free RNA derived from a blood sample from the pregnant mother. The methods can detect deletions or duplications of fetal cells, cell-free DNA, or cell-free RNA even if the amount of cells, cell-free DNA, or cell-free RNA from the pregnant mother is large. In some embodiments, invasive prenatal testing involves testing DNA or RNA derived from a fetal sample (e.g., CVS or amniocentesis sample). If the sample is contaminated with DNA or RNA from the pregnant mother, the methods can be used to detect deletions or duplications of fetal DNA or RNA.
[0240] In addition to determining the presence or absence of copy number diversity, one or more other factors can be analyzed as needed. These factors can be used to improve the accuracy of diagnosis (e.g., determining the presence or absence of cancer or an increased risk of cancer, cancer classification, or cancer staging) or prognosis. These factors can also be used to select specific treatments or treatment plans that are likely to be effective for the patient. Examples of such factors include the presence or absence of polymorphisms or mutations, the total amount of cell-free DNA, cell-free RNA, or any specific change (increase or decrease) in microRNA (miRNA), altered (increased or decreased) tumor fractions, changes (increases or decreases) in methylation, altered (increased or decreased) DNA integrity, altered (increased or decreased) or alternative mRNA splicing, etc.
[0241] The following describes methods for detecting deletions or duplications using phase-determined data (e.g., estimated or measured phase-determined data) or unphase-determined data, testable samples, methods for preparing, amplifying, and quantifying samples, phase-determining methods for genetic data, detectable polymorphisms, mutations, nucleic acid changes, mRNA splicing changes, and nucleic acid quantity changes, databases of results derived from the said methods, other risk factors and screening methods, diagnosable or treatable cancers, cancer treatments, cancer models for testing therapies, and methods for prescribing and implementing therapies.
[0242] Example of a method for determining ploidy using phase-determined data Some of the methods described above in the present invention are based in part on the finding that the false negative and false positive rates are reduced when phase-determined data is used for CNV detection compared to when phase-determined data is used (Figures 20A to 27). This improvement is particularly beneficial for samples containing CNVs in small amounts. Therefore, phase-determined data improves the accuracy of CNV detection compared to using phase-determined data (for example, a method for calculating aggregated allele ratios to give the ratio or aggregate value (e.g., mean) of alleles present at one or more loci of a chromosome or chromosomal segment, without considering that allele ratios at different loci may indicate the presence of the same or different haplotypes in abnormal amounts). By using phase-determined data, it is possible to more accurately determine whether the difference between the measured allele ratio and the expected allele ratio is due to noise or the presence of CNVs. For example, if the difference between the measured allele ratio and the expected allele ratio at most or all loci in a region indicates an overpopulation of the same haplotype, the presence of CNVs is more likely. Using allele linkage for a given haplotype, it is possible to determine whether measured genetic data is consistent with the overpopulation of the same haplotype (rather than random noise). In contrast, when the difference between the measured allele ratio and the expected allele ratio is due solely to noise (e.g., experimental error), in some cases, the first haplotype appears to be overpopulated for about half the time, and the second haplotype appears to be overpopulated for the remaining half.
[0243] Accuracy can be improved by considering the likelihood of linkage between SNPs and crossover that occurred during meiosis, which produced gametes that formed embryos that developed into fetuses. Using linkage when constructing the expected distribution of alleles for one or more hypotheses results in a more accurate representation of the actual alleles compared to not using linkage. For example, suppose there are two closely spaced SNPs, S1 and S2, where SNP1 is A and SNP2 is A in one homolog 1 of the mother, and SNP1 is B and SNP2 is B in the other homolog 2. If both SNPs in the father's homologs are A, and B is measured at SNP1 in the fetus, it indicates that homolog 2 was inherited by the fetus, and therefore the likelihood of B being present at SNP2 in the fetus is higher. Models that consider linkage can predict this, but models that do not consider linkage cannot. Alternatively, if the mother's SNP1 is AB and the neighboring SNP2 is AB, two hypotheses can be used that correspond to the mother's triplicity at this location: the hypothesis that it involves a congruent copy error (nondisjunction during meiosis II or mitosis in early fetal development) and the hypothesis that it involves a partially congruent copy error (nondisjunction during meiosis I). If it is a congruent copy error, then if the fetus inherited AA from the mother for SNP1, it is more likely that the fetus inherited either AA or BB for SNP2 rather than AB. If it is a partially congruent copy error, then the fetus inherited AB from the mother for both SNP1 and SNP2. Allele distribution hypotheses created with linkage-considered CNV calling can make these predictions. Therefore, they can better correspond to actual allele measurements than CNV calling without linkage.
[0244] In some embodiments, phase-determined gene data is used to determine whether there is copy number overpopulation of the first homologous chromosome segment compared to the second homologous chromosome segment in the individual's genome (e.g., in the genome of one or more cells, or in cell-free DNA or cell-free RNA). Examples of overpopulation include duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, there is no overpopulation because the first and second homologous chromosome segments are present in equal proportions (e.g., one copy of each segment in a diploid sample). In some embodiments, the calculated allele ratio of a nucleic acid sample is compared to the expected allele ratio to determine whether there is overpopulation, as further described below. In this specification, the term “first homologous chromosome segment compared to the second homologous chromosome segment” means the first homolog of a given chromosome segment and its second homolog.
[0245] In some embodiments, the method includes obtaining phase-determined genetic data of the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci within the first homologous chromosome segment; obtaining phase-determined genetic data of the second homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci within the second homologous chromosome segment; and obtaining measured allele genetic data for each allele present at each locus in the set of polymorphic loci, including the amount of each allele present in DNA or RNA samples derived from one or more target cells and one or more non-target cells from the individual. In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of a first homologous chromosome segment; for each of the hypotheses, calculating expected gene data for multiple loci in the sample from the obtained phase-determined gene data for one or more possible ratios of DNA or RNA derived from the one or more target cells to the total DNA or RNA in the sample; calculating the data fit (by computer, etc.) between the obtained sample gene data and the expected gene data for the sample for each possible ratio of DNA or RNA and for each hypothesis; ranking one or more hypotheses according to the data fit; selecting the highest-ranked hypothesis to determine the degree of overpopulation of the copy number of a first homologous chromosome segment in the genome of one or more cells derived from the individual.
[0246] In one aspect, the present invention features a method for determining the copy number of a target chromosome or chromosome segment in the genome of a fetus. In some embodiments, the method comprises obtaining phase-determined genetic data of at least one biological parent of the fetus, wherein the phase-determined genetic data includes the identity of alleles present at each locus of a set of polymorphic loci in a first homologous chromosome segment and a second homologous chromosome segment of the parent. In some embodiments, the method comprises obtaining genetic data for a set of polymorphic loci in a chromosome or chromosome segment in a mixed sample of DNA or RNA, including fetal DNA or RNA and maternal DNA or RNA derived from the mother of the fetus, by measuring the amount of each allele present at each locus. In some embodiments, the method comprises listing one or more sets of hypotheses that define the copy number of the target chromosome or chromosome segment present in the genome of the fetus. In some embodiments, the method includes, for each of the hypotheses, creating (i) a probability distribution of the expected amount of each allele present at each of the multiple loci in the mixed sample (for example, on a computer) from the obtained phase-determined gene data from the (both) parents and optionally (ii) the probability of one or more crossovers that may have occurred during the formation of the gamete that resulted in a copy of the target chromosome or chromosome segment in the fetus; for each of the hypotheses, calculating (i) the fit between the obtained gene data of the mixed sample and (ii) the probability distribution of the expected amount of each allele at each of the multiple loci in the mixed sample (for each hypothesis) (on a computer, etc.); ranking one or more hypotheses according to the data fit; selecting the highest-ranked hypothesis; and determining the copy number of the target chromosome segment in the fetal genome.
[0247] In some embodiments, the method includes obtaining phase-determined genetic data using any of the methods disclosed herein or any well-known method. In some embodiments, the method includes, simultaneously or sequentially, (i) obtaining phase-determined genetic data for the first homologous chromosome segment, including the identity of the allele present at each locus in the set of polymorphic loci of the first homologous chromosome segment; (ii) obtaining phase-determined genetic data for the second homologous chromosome segment, including the identity of the allele present at each locus in the second homologous chromosome segment, for each locus in the set of polymorphic loci of the second homologous chromosome segment; and (iii) obtaining measured allele genetic data, including the amount of each allele present at each locus in the set of polymorphic loci in a DNA sample derived from one or more cells from the individual.
[0248] In some embodiments, the method includes calculating allele ratios for one or more loci from a set of polymorphic loci that are heterozygous in at least one cell from which the sample originates (e.g., loci that are heterozygous in the fetus and / or in the mother). In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele by the total measurement of all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele (e.g., an allele in the first homologous chromosome segment) by the measurement of one or more other alleles at that locus (e.g., an allele in the second homologous chromosome segment). The calculated allele ratios may be calculated using any of the methods disclosed herein or by any standard method (e.g., any mathematical transformation of the calculated allele ratios described herein).
[0249] In some embodiments, the method includes determining whether there is a copy number excess of the first homologous chromosome segment by comparing one or more calculated allele ratios of a given locus with the allele ratio expected for that locus if the first and second homologous chromosome segments are present in equal proportions. In some embodiments, the expected allele ratio assumes that all possible alleles at a given locus have an equal likelihood of existence. In some embodiments, where the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele by the total measurement of all alleles at that locus, the corresponding expected allele ratio is 0.5 for bialleletic loci and 1 / 3 for trialleletic loci. In some embodiments, the expected allele ratio is the same for all loci, such as 0.5. In some embodiments, the expected allele ratio assumes that all possible alleles at a given locus may have different likelihoods of existence (e.g., likelihoods based on the respective frequencies of the allele in a particular population to which the subject belongs (e.g., a population based on the subject's ancestry)). Such allele frequencies are publicly available (see, for example, the HapMap Project, the Perlegen Human Haplotype Project, the webpage (ncbi.nlm.nih.gov / projects / SNP / ), Sherry ST, Ward MH, Kholodov M, et al. dbSNP, the NCBI Genetic Variation Database, Nucleic Acids Res. 2001 Jan 1;29(1):308-11, each incorporated herein by reference). In some embodiments, the expected allele ratio is the expected allele ratio for a particular individual being tested for a particular hypothesis defining the degree of overpopulation of the first homologous chromosome segment. For example, the expected allele ratio for a particular individual may be determined based on phase-determined or phase-undetermined genetic data from the individual (e.g., a sample from an individual that is thought to be free of deletions or duplications, such as a non-cancerous sample), or data from one or more close relatives of the individual.In some embodiments, in the case of prenatal testing, the expected allele ratio is the expected allele ratio for a mixed sample (e.g., maternal plasma, or a serum sample containing maternal cell-free DNA and fetal cell-free DNA) containing DNA or RNA from the pregnant mother and fetus, for a specific hypothesis defining the degree of excess occurrence of the first homologous chromosome segment. For example, the expected allele ratio of the mixed sample may be determined based on maternal genetic data and predicted genetic data for the fetus (e.g., predictions about alleles likely to be inherited from the mother and / or father). In some embodiments, phase-determined or phase-undetermined genetic data from a maternal-only DNA or RNA sample (e.g., pia mater from a maternal blood sample) is used to determine alleles from maternal DNA or RNA in the mixed sample, and alleles that the fetus may have inherited from the mother (and therefore may be present in the fetal DNA or RNA in the mixed sample). In some embodiments, phase-determined or phase-undetermined genetic data derived from a DNA or RNA sample originating solely from the father are used to determine alleles that the fetus may have inherited from the father (and therefore may be present in the fetal DNA or RNA in the mixed sample). The expected allele ratios may be calculated using any of the methods disclosed herein or any standard method (e.g., any mathematical transformation of the expected allele ratios described herein) (U.S. Patent Publication No. 2012 / 0270212, filed November 18, 2011, which is incorporated herein by reference in its entirety).
[0250] In some embodiments, the calculated allele ratio indicates copy number overpopulation of the first homologous chromosome segment if (i) the allele ratio obtained by dividing the measured amount of alleles present at that locus on the first homologous chromosome by the total measured amount of all alleles at that locus is greater than the expected allele ratio for that locus, or (ii) the allele ratio obtained by dividing the measured amount of alleles present at that locus on the second homologous chromosome by the total measured amount of all alleles at that locus is less than the expected allele ratio for that locus. In some embodiments, the calculated allele ratio is considered to indicate overpopulation only if it is significantly greater than or less than the expected ratio for that locus. In some embodiments, the calculated allele ratio indicates no copy number overpopulation of the first homologous chromosome segment if (i) the allele ratio obtained by dividing the measured amount of alleles present at that locus on the first homologous chromosome by the total measured amount of all alleles at that locus is less than or equal to the expected allele ratio of that locus, or (ii) the allele ratio obtained by dividing the measured amount of alleles present at that locus on the second homologous chromosome by the total measured amount of all alleles at that locus is greater than or equal to the expected allele ratio of that locus. In some embodiments, the calculated ratio equal to the corresponding expected ratio is ignored (to indicate no overpopulation).
[0251] In various embodiments, one or more calculated allele ratios are compared to the corresponding expected allele ratio using one or more of the following methods. In some embodiments, for a particular locus, it is determined whether the calculated allele ratio is greater than or less than the expected allele ratio, regardless of the magnitude of the difference. In some embodiments, for a particular locus, the magnitude of the difference between the calculated allele ratio and the expected allele ratio is determined, regardless of whether the calculated allele ratio is greater than or less than the expected allele ratio. In some embodiments, for a particular locus, it is determined whether the calculated allele ratio is greater than or less than the expected allele ratio and the magnitude of the difference. In some embodiments, regardless of the magnitude of the difference, it is determined whether the mean or weighted mean of the calculated allele ratios is greater than or less than the mean or weighted mean of the expected allele ratios. In some embodiments, regardless of whether the mean or weighted mean of the calculated allele ratios is greater than or less than the mean or weighted mean of the expected allele ratios, the magnitude of the difference between the mean or weighted mean of the calculated allele ratios and the mean or weighted mean of the expected allele ratios is determined. In some embodiments, the mean or weighted mean of the calculated allele ratios is determined to be greater than or less than the mean or weighted mean of the expected allele ratios, and the magnitude of the difference is determined. In some embodiments, the mean or weighted mean of the magnitude of the difference between the calculated allele ratios and the expected allele ratios is determined.
[0252] In some embodiments, the magnitude of the difference between the calculated allele ratio and the expected allele ratio for one or more loci is used to determine whether the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment in one or more genomes of the cell.
[0253] In some embodiments, copy number overpopulation of the first homologous chromosome segment is determined if one or more of the following conditions are met: In some embodiments, the number of calculated allele ratios indicating copy number overpopulation of the first homologous chromosome segment is greater than a threshold. In some embodiments, the number of calculated allele ratios indicating that the first homologous chromosome segment is not copy number overpopulation is less than a threshold. In some embodiments, the magnitude of the difference between the calculated allele ratio indicating copy number overpopulation of the first homologous chromosome segment and the corresponding expected allele ratio is greater than a threshold. In some embodiments, the sum of the magnitudes of the differences between all calculated allele ratios indicating overpopulation and the corresponding expected allele ratios is greater than a threshold. In some embodiments, the magnitude of the difference between the calculated allele ratio indicating that the first homologous chromosome segment is not copy number overpopulation and the corresponding expected allele ratio is less than a threshold. In some embodiments, the average or weighted average of the calculated allele ratios obtained by dividing the measured amount of alleles present on the first homologous chromosome by the total measured amount of all alleles at that locus is greater than the average or weighted average of the expected allele ratios divided by at least a threshold. In some embodiments, the average or weighted average of the calculated allele ratios obtained by dividing the measured amount of alleles present on the second homologous chromosome by the total measured amount of all alleles at that locus is less than the average or weighted average of the expected allele ratios divided by at least a threshold. In some embodiments, the data fit between the calculated allele ratios and the allele ratios predicting copy number overpopulation of the first homologous chromosome segment is less than a threshold (indicating good data fit). In some embodiments, the data fit between the calculated allele ratios and the allele ratios indicating that there is no copy number overpopulation of the first homologous chromosome segment is greater than a threshold (indicating poor data fit).
[0254] In some embodiments, copy number overpopulation of the first homologous chromosome segment is determined if one or more of the following conditions are met: In some embodiments, the number of calculated allele ratios indicating copy number overpopulation of the first homologous chromosome segment is less than a threshold. In some embodiments, the number of calculated allele ratios indicating that the first homologous chromosome segment is not copy number overpopulation is greater than a threshold. In some embodiments, the magnitude of the difference between the calculated allele ratio indicating copy number overpopulation of the first homologous chromosome segment and the corresponding expected allele ratio is less than a threshold. In some embodiments, the magnitude of the difference between the calculated allele ratio indicating that the first homologous chromosome segment is not copy number overpopulation and the corresponding expected allele ratio is greater than a threshold. In some embodiments, the difference between the average or weighted average of the calculated allele ratios obtained by dividing the measured amount of alleles present on the first homologous chromosome by the total measured amount of all alleles at that locus, and the average or weighted average of the expected allele ratio is less than a threshold. In some embodiments, the difference between the mean or weighted mean of the expected allele ratio and the mean or weighted mean of the calculated allele ratio, obtained by dividing the measured amount of alleles present on the second homologous chromosome by the total measured amount of all alleles at that locus, is less than the threshold. In some embodiments, the data fit between the calculated allele ratio and the allele ratio predicting copy number overpopulation of the first homologous chromosome segment is greater than the threshold. In some embodiments, the data fit between the calculated allele ratio and the allele ratio indicating that there is no copy number overpopulation of the first homologous chromosome segment is less than the threshold. In some embodiments, the threshold is determined by testing samples known to have the CNV of the target and / or samples known to be without the CNV.
[0255] In some embodiments, determining whether there is copy number overpopulation of the first homologous chromosome segment involves enumerating a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment. One example of a hypothesis is that there is no overpopulation if the first and second homologous chromosome segments are present in equal proportions (e.g., one copy of each segment in a diploid sample). Another example of a hypothesis involves a first homologous chromosome segment that is duplicated one or more times (e.g., having one, two, three, four, five or more extra copies of the first homologous chromosome compared to the copy number of the second homologous chromosome segment). Yet another example of a hypothesis involves deletion of the second homologous chromosome segment. Another exemplary hypothesis is deletion of both the first and second homologous chromosome segments. In some embodiments, the predicted allele ratios for loci that are heterozygous in at least one cell (e.g., the loci that are heterozygous in the fetus and / or the loci that are heterozygous in the mother) are estimated for each hypothesis, assuming that the excess occurrence is defined by the hypothesis. In some embodiments, the likelihood of a hypothesis being correct is calculated by comparing the calculated allele ratios with the predicted allele ratios, and the most likely hypothesis is selected.
[0256] In some embodiments, the expected distribution of the test statistic is calculated for each hypothesis using the predicted allele ratio. In some embodiments, the expected distribution of the test statistic calculated using the calculated allele ratio is compared with the expected distribution of the test statistic calculated using the predicted allele ratio to calculate the likelihood that the hypothesis is correct, and the most likely hypothesis is selected.
[0257] In some embodiments, the predicted allele ratio for a locus that is heterozygous in at least one cell (e.g., the locus that is heterozygous in the fetus and / or the locus that is heterozygous in the mother) is estimated based on the phase-determined gene data of the first homologous chromosome segment, the phase-determined gene data of the second homologous chromosome segment, and the assumption that the excess frequency is defined by a hypothesis. In some embodiments, the calculated allele ratio is compared with the predicted allele ratio to calculate the likelihood that the hypothesis is correct, and the most likely hypothesis is selected.
[0258] Use of mixed samples In many embodiments, the sample is a mixed sample containing DNA or RNA derived from one or more target cells and one or more non-target cells. In some embodiments, the target cells are cells having CNVs such as the target deletion or duplication, and the non-target cells are cells that do not have the target copy number diversity (e.g., a mixture of cells having the target deletion or duplication and cells that do not have either the target deletion or duplication). In some embodiments, the target cells are cells associated with disease or disability or an increased risk of disease or disability (e.g., cancer cells), and the non-target cells are cells not associated with disease or disability or an increased risk of disease or disability (e.g., non-cancerous cells). In some embodiments, all of the target cells have the same CNV. In some embodiments, two or more target cells have different CNVs. In some embodiments, one or more of the target cells have a CNV, polymorphism, or mutation associated with disease or disability or an increased risk of disease or disability that is not found in at least one other target cell. In some such embodiments, it is assumed that the fraction of cells from a given sample that are associated with disease or disability or an increased risk of disease or disability is equal to or greater than the fraction with the highest frequency of these CNVs, polymorphisms, or mutations in that sample. For example, if 6% of the cells have K-ras mutations and 8% have BRAF mutations, it is assumed that at least 8% of the cells are cancerous.
[0259] In some embodiments, the ratio of DNA (or RNA) derived from one or more target cells in the sample to the total DNA (or RNA) is calculated. In some embodiments, a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment is enumerated. In some embodiments, the predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother) is estimated for each hypothesis, assuming that the calculated ratio of DNA or RNA and the degree of overpopulation are defined by the hypothesis. In some embodiments, the likelihood of a hypothesis being correct is calculated by comparing the calculated allele ratio with the predicted allele ratio, and the most likely hypothesis is selected.
[0260] In some embodiments, the expected distribution of the test statistic calculated using the predicted allele ratios and the calculated ratios of DNA or RNA is estimated for each hypothesis. In some embodiments, the likelihood of a hypothesis being correct is determined by comparing the test statistic calculated using the calculated allele ratios and the calculated ratios of DNA or RNA with the expected distribution of the test statistic calculated using the predicted allele ratios and the calculated ratios of DNA or RNA, and the hypothesis with the highest probability is selected.
[0261] In some embodiments, the method includes listing a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment. In some embodiments, for each hypothesis, the method includes estimating either (i) a predicted allele ratio for a locus that is heterozygous in at least one cell (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother) given the degree of overpopulation defined by the hypothesis, or (ii) an expected distribution of a test statistic calculated using the predicted allele ratio and the expected ratio of the DNA or RNA derived from the one or more target cells to the total DNA or RNA in the sample for one or more possible ratios of DNA or RNA. In some embodiments, data fit is calculated by comparing either (i) the calculated allele ratio to the predicted allele ratio, or (ii) a test statistic calculated using the calculated allele ratio and the expected distribution of the test statistic calculated using the predicted allele ratio and the expected ratio of DNA or RNA, with the expected distribution of the test statistic calculated using the predicted allele ratio and the expected ratio of DNA or RNA. In some embodiments, one or more hypotheses are ranked according to the data fit, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm (e.g., a search algorithm) is used in one or more of the steps of calculating the data fit, ranking the hypotheses, and selecting the highest-ranked hypothesis. In some embodiments, the data fit is a fit to a beta-binomial distribution or a fit to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, maximum posterior probability estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained gene data and the expected gene data.
[0262] In some embodiments, the method includes creating a range of possible ratios from a lower limit to an upper limit for the ratio of DNA or RNA derived from the one or more target cells to the total DNA or RNA in the sample. In some embodiments, a set of one or more hypotheses defining the degree of overpopulation of the first homologous chromosome segment is enumerated. In some embodiments, for each possible ratio of DNA or RNA in the range and each hypothesis, the method includes estimating either (i) a predicted allele ratio of a locus that is heterozygous in at least one cell (e.g., a locus that is heterozygous in the fetus and / or a locus that is heterozygous in the mother) given the possible ratio of DNA or RNA and the degree of overpopulation defined by the hypothesis, or (ii) an expected distribution of a test statistic calculated using the predicted allele ratio and the possible ratio of DNA or RNA. In some embodiments, the method includes calculating the likelihood that a hypothesis is correct for each ratio of the DNA or RNA in a given category and each hypothesis by comparing (i) a calculated allele ratio to the predicted allele ratio or (ii) a test statistic calculated using the calculated allele ratio and the possible ratio of the DNA or RNA with the expected distribution of the test statistic calculated using the predicted allele ratio and the possible ratio of the DNA or RNA. In some embodiments, a composite probability of each hypothesis is determined by combining the probabilities of the hypotheses for each ratio of possible ratios in the given category, and the hypothesis with the highest composite probability is selected. In some embodiments, the composite probability of each hypothesis is determined by taking a weighted average of the hypothesis probabilities for a particular possible ratio based on the likelihood that the possible ratio is the correct ratio.
[0263] In some embodiments, the ratio of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample is estimated using a technique selected from the group consisting of maximum likelihood estimation, maximum posterior probability estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, it is assumed that the ratio of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample is the same for two or more (or all) of the target CNVs. In some embodiments, the ratio of DNA or RNA derived from one or more target cells to the total DNA or RNA in the sample is calculated for each target CNV. Example of a method using insufficient phase-determined data
[0264] Naturally, in many aspects, insufficient phase-determined data is used. For example, it may not be 100% certain which alleles are present at one or more loci in the first homologous chromosome segment and / or the second homologous chromosome segment. In some aspects, the priority of possible haplotypes of an individual (e.g., haplotypes based on population-based haplotype frequencies) is used to calculate the probability of each hypothesis. In some aspects, the priority of possible haplotypes is adjusted using phase-determined data from other methods of phase-determining genetic data or from other subjects (e.g., high-priority subjects) to improve the population data used in informatics based on the phase determination of the individual.
[0265] In some embodiments, the phase-determined gene data includes probability data relating to two or more possible sets of phase-determined gene data, each of which includes possible identities of alleles present at each locus in the set of polymorphic loci in the first homologous chromosome segment and possible identities of alleles present at each locus in the set of polymorphic loci in the second homologous chromosome segment. In some embodiments, the probability of at least one hypothesis is determined for each of the possible sets of phase-determined gene data. In some embodiments, the composite probability of the hypothesis is determined by combining the probabilities of the hypotheses for each of the possible sets of phase-determined gene data, and the hypothesis with the highest composite probability is selected.
[0266] Insufficient phase-determined data for use in the methods of the present invention may be prepared using any of the methods disclosed herein or any known method (e.g., inferring the most likely phase using population-based haplotype frequencies). In some embodiments, phase-determined data can be obtained by probabilistically combining haplotypes of smaller segments. For example, possible haplotypes can be determined based on possible combinations of haplotypes in other regions of the same chromosome as a haplotype in a first region. The probability that a particular haplotype in a different region is part of the same, larger haplotype block on the same chromosome can be determined, for example, using population-based haplotype frequencies and / or known recombination rates between different regions.
[0267] In some embodiments, a single-hypothesis rejection test is used for the null hypothesis of disomality. In some embodiments, the probability of the disomality hypothesis is calculated, and if this probability is below a predetermined threshold (e.g., less than 1 in 1000), the disomality hypothesis is rejected. If the null hypothesis is rejected, it may be due to errors in insufficient phase-determined data or the presence of CNVs. In some embodiments, more accurate phase-determined data is available (e.g., phase-determined data obtained from one of the molecular phase-determining methods disclosed herein to obtain actual phase-determined data, rather than phase-determined data inferred based on bioinformatics). In some embodiments, the probability of the disomality hypothesis is recalculated using the aforementioned more accurate phase-determined data to determine whether the disomality hypothesis is still rejected. This rejection of the hypothesis indicates the presence of the duplication or deletion of the chromosomal segment. If necessary, the false positive rate can be changed by adjusting the threshold.
[0268] Another example of an embodiment in which ploidy is determined using phase-determined data. In exemplary embodiments of this specification, a method for determining the ploidy of chromosome segments in an individual sample is provided. The method includes the following steps: a. Allele frequency data is received, which includes the amount of each allele present in each locus among the set of polymorphic loci in the chromosome segment in the sample. b. By estimating the phase of the allele frequency data, phase-determined allele information relating to the set of polymorphic loci is created. c. Using the allele frequency data, create individual probabilities of the allele frequencies of the polymorphic loci for different ploidy states. d. Using the individual probabilities and the phase-determined allele information, create a joint probability for the set of polymorphic loci. e. Determining the ploidy of the chromosome segment by selecting the optimal model that exhibits chromosomal ploidy based on the joint probability described above.
[0269] As disclosed herein, allele frequency data (also referred to herein as genetic data of the measured alleles) can be prepared by methods known in the art. For example, such data can be prepared using quantitative PCR or microarrays. In one exemplary embodiment, such data is prepared using nucleic acid sequence data, particularly high-processing nucleic acid sequence data.
[0270] In one specific example, errors in the allele frequency data are corrected before individual probabilities are created using the allele frequency data. In a particular exemplary embodiment, the corrected errors include bias in allele amplification efficiency. In other embodiments, the corrected errors include environmental impurities and impurities from genotyping analysis. In some embodiments, the corrected errors include bias from allele amplification, environmental impurities, and impurities from genotyping analysis.
[0271] In certain embodiments, the individual probabilities are constructed using a set of models for both different ploidy states and allele imbalance rates relating to the set of polymorphic loci. In these embodiments and other embodiments, the joint probabilities are constructed by considering linkages between polymorphic loci in the chromosome segments.
[0272] Accordingly, an exemplary embodiment of this specification, combining some of these embodiments, provides a method for detecting chromosome ploidy in a sample of an individual. The method comprises the following steps: a. Receiving nucleic acid sequence data relating to alleles present in a set of polymorphic loci in the chromosomal segments of the aforementioned individual, b. Using the nucleic acid sequence data, detect the allele frequency in the set of polymorphic loci, c. Correct the bias in the allele amplification efficiency of the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci, d. By estimating the phase of the nucleic acid sequence data, phase-determined allele information relating to the set of polymorphic loci is created. e. By comparing the modified allele frequencies with a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci, we can create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states. f. By combining the individual probabilities while considering the linkage between polymorphic loci in the chromosome segment, a joint probability relating to the set of polymorphic loci is created. g. Select the optimal model that exhibits chromosomal aneuploidy based on the aforementioned joint probability.
[0273] As disclosed herein, the individual probabilities can be generated using a set of models or hypotheses for both different ploidy states and mean allele misbalance rates relating to the set of polymorphic loci. For example, in a particular case, the individual probabilities are generated by modeling the ploidy states of the first homolog of the chromosome segment and the second homolog of the chromosome segment. The ploidy states to be modeled include:
[0274] (1) All cells are free from deletions or amplifications in the first or second homolog of the chromosome segment.
[0275] (2) At least some cells have a deletion of the first homolog of the chromosome segment or an amplification of the second homolog.
[0276] (3) At least some cells have a deletion of the second homolog of the chromosome segment or an amplification of the first homolog.
[0277] Furthermore, the above models can also be described as hypotheses used to suppress certain models. Therefore, the three hypotheses shown above are those that can be used.
[0278] The modeled mean allele disequilibrium rate includes any range of mean allele disequilibrium rates, including the actual mean allele disequilibrium rates of the chromosome segments. For example, in certain exemplary embodiments, the range of modeled mean allele disequilibrium rates may be lower bounds of 0, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5%, and upper bounds of 1, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, and 99%. The modeling interval for this range may be any interval depending on the computing power used and the time available for analysis. For example, the interval can be modeled as 0.01, 0.05, 0.02, or 0.1.
[0279] In certain exemplary embodiments, the sample has an average allele-to-chromosome mismatch rate of 0.4% to 5%. In certain embodiments, the average allele-to-chromosome mismatch rate is small. In these embodiments, the average allele-to-chromosome mismatch rate is typically less than 10%. In certain exemplary embodiments, the allele-to-chromosome mismatch rate has a lower limit of 0.25, 0.3, 0.4, 0.5, 0.6, 0.75, 1, 2, 2.5, 3, 4, and 5%, and an upper limit of 1, 2, 2.5, 3, 4, and 5%. In other exemplary embodiments, the average allele-to-chromosome mismatch rate has a lower limit of 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0%, and an upper limit of 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 3.0, 4.0, or 5.0%. For example, in one case, the average allele mismatch rate of the sample is 0.45 to 2.5%. In other cases, the average allele mismatch rate is detected with a sensitivity of 0.45, 0.5, 0.6, 0.8, 0.8, 0.9, or 1.0. Examples of low allele mismatch rate samples used in the method of the present invention include plasma samples derived from individuals with cancer containing circulating tumor DNA, or plasma samples derived from pregnant women containing circulating fetal DNA.
[0280] In the case of SNVs, the proportion of abnormal DNA is usually measured using the mutant allele frequency (number of mutant alleles present at a given locus / total number of alleles present at that locus). Since the difference in the amount of two homologs in a tumor is similar, the proportion of abnormal DNA in CNVs is measured using the mean allele-allelic ratio (AAI). The mean allele-allelic ratio is defined as |(H1-H2)| / (H1+H2), where Hi is the average copy number of homolog i in the sample, and Hi / (H1+H2) is the fractional amount of homolog i, i.e., the homolog ratio. The maximum homolog ratio is the homolog ratio of the homolog with the highest quantity.
[0281] The trial dropout rate is the percentage of read-less SNPs estimated using all SNPs. The single-allelic dropout (ADO) rate is the percentage of SNPs with only one allele estimated using only heterozygous SNPs. Genotype confidence can be determined by fitting a binomial distribution to the number of reads present in each SNP that was a B allele read, and then estimating the probability of each genotype using the ploidy status of the region of interest of that SNP.
[0282] In tumor tissue samples, chromosomal aneuploidy (exemplified in this paragraph as CNV) can be indicated by transitions between allele frequency distributions. In plasma samples, CNVs can be identified by a maximum likelihood algorithm that searches for plasma CNVs in regions where tumor samples from the same individual also have CNVs, using haplotype information estimated from the tumor sample. This algorithm can model expected allele frequencies at 0.025% intervals across all allele imbalance rates for three sets of hypotheses: (1) all cells are normal (no allele imbalance), (2) some / all cells have homolog 1 deletion or homolog 2 amplification, or (3) some / all cells have homolog 2 deletion or homolog 1 amplification. In certain exemplary embodiments considering SNP locus linkage, as illustrated herein, the likelihood of each hypothesis can be determined for each SNP using a Bayesian classifier based on a beta-binomial model of expected and observed allele frequencies for all heterozygous SNPs, and then the simultaneous likelihood of multiple SNPs can be calculated. Next, we can select the most likely hypothesis.
[0283] Consider a chromosomal region of a tumor that has an average of N copies. Here, c represents the DNA fraction in plasma derived from a mixture of normal and tumor cells in a region of two chromosomes. The AAI is calculated as follows:
number
[0284] In certain examples, errors in the allele frequency data are corrected before individual probabilities are created using the allele frequency data. Different types of error and / or bias corrections are disclosed herein. In certain exemplary embodiments, the corrected error is a bias in allele amplification efficiency. In other embodiments, the corrected error includes environmental impurities and impurities from genotyping analysis. In some embodiments, the corrected error includes bias from allele amplification, environmental impurities, and impurities from genotyping analysis.
[0285] Naturally, bias in allele amplification efficiency can be determined for an allele as part of an experimental determination involving the test sample, or it can be determined at different points in time using a sample group containing alleles for which efficiency has been calculated. Environmental impurities and impurities determined by genotyping analysis are usually determined in the same single procedure as the analysis of the test sample.
[0286] In certain embodiments, environmental impurities and impurities determined by genotypic analysis are determined for homozygous alleles in the sample. Naturally, even if a locus is selected for analysis because it exhibits relatively high heterozygosity within a population, some loci in any sample originating from a given individual will be heterozygous, while others will be homozygous. In some embodiments, the polyploidy of a chromosome segment may be determined using heterozygous loci of an individual, but it is beneficial to use homozygous loci to calculate environmental impurities and impurities determined by genotypic analysis.
[0287] In certain cases, the selection is made by analyzing the magnitude of the difference between the phase-determined allele information and the estimated allele frequency with respect to the model.
[0288] In some cases, the individual probabilities of allele frequencies are constructed based on a beta-binomial model of the expected and observed allele frequencies for the set of polymorphic loci. In other cases, the individual probabilities are constructed using a Bayesian classifier.
[0289] In certain exemplary embodiments, the nucleic acid sequence data is created by performing high-process DNA sequencing on multiple copies of a series of amplification products generated by a multiple amplification reaction, where each amplification product of the series is amplified across at least one polymorphic locus of the set of polymorphic loci, and in certain embodiments where each polymorphic locus of the set of polymorphic loci is amplified, the multiple amplification reaction is performed with restrictive primer conditions for at least half of the reaction. In some embodiments, restrictive primer concentrations are used for 1 / 10, 1 / 5, 1 / 4, 1 / 3, 1 / 2, or all of the multiple reactions. Factors that are likely to yield restrictive primer conditions in amplification reactions, such as PCR, are provided herein.
[0290] In certain embodiments, the methods provided herein detect the ploidy of multiple chromosome segments across multiple chromosomes. Thus, chromosome ploidy in these embodiments is determined for a set of chromosome segments in the sample. These embodiments require a greater number of multiple amplification reactions. Therefore, the multiple amplification reactions in these embodiments may include, for example, 2,500 to 50,000 multiple reactions. In certain embodiments, the multiple reactions are performed in the range of 100, 200, 250, 500, 1,000, 2,500, 5,000, 10,000, 20,000, 25,000, and 50,000 with a lower limit of 100,250, 500, 1,000, 2,500, 5,000, 10,000, 20,000, 25,000, 50,000, and 100,000.
[0291] In exemplary embodiments, the set of polymorphic loci is a set of loci known to exhibit high heterozygosity. However, some of these loci in any given individual are expected to be homozygous. In specific exemplary embodiments, the method of the present invention utilizes nucleic acid sequence information from both homozygous and heterozygous loci of an individual. For example, the homozygous loci of an individual are used for error correction, and the heterozygous loci are used to determine the allele imbalance rate of a sample. In specific embodiments, at least 10% of the polymorphic loci are heterozygous loci of the individual.
[0292] As disclosed herein, the analysis of target SNP loci known to be heterozygous in the population is preferred. Accordingly, in certain embodiments, polymorphic loci are selected in which at least 10, 20, 25, 50, 75, 80, 90, 95, 99, or 100% are known to be heterozygous in the population.
[0293] As disclosed herein, in certain embodiments, the sample is a plasma sample derived from a pregnant woman.
[0294] In some examples, the method further comprises performing the method on a control sample having a known mean allele disequilibrium ratio. The control can reproduce the mean allele disequilibrium ratio of alleles in the sample that are present at low concentrations, such as those predicted for circulating free DNA derived from a fetus or tumor, with a mean allele disequilibrium ratio of 0.4–10% for a specific allele state exhibiting aneuploidy of the chromosomal segment.
[0295] As disclosed herein, in some embodiments, a PlasmArt control is used as the control. In one aspect, the control is a sample prepared by a method comprising fragmenting a nucleic acid sample known to exhibit chromosomal aneuploidy to the size of DNA fragments circulating in the plasma of the individual. In another aspect, a control without aneuploidy in chromosomal segmentation is used.
[0296] In exemplary embodiments, data from one or more controls can be analyzed together with the test sample using the method described above. Examples of controls include another sample from an individual not suspected of having chromosomal aneuploidy, or a sample suspected of containing CNV or chromosomal aneuploidy. For example, if the test sample is a plasma sample suspected of containing circulating free tumor DNA, the method can be performed on a control sample derived from the tumor of the subject together with this plasma sample. As disclosed herein, the control sample can be prepared by fragmenting a DNA sample known to exhibit chromosomal aneuploidy. In particular, if the sample is from an individual suffering from cancer, such fragmentation can yield a DNA sample that replicates the DNA composition of apoptotic cells. Data from the control sample enhance the reliability of detecting chromosomal aneuploidy.
[0297] In certain embodiments of the method for determining ploidy, the sample is a plasma sample derived from an individual suspected of having cancer. In these embodiments, the method further includes determining, based on the selection, whether copy number diversity exists in the tumor cells of the individual. In these embodiments, the sample may be a plasma sample derived from a certain individual. In these embodiments, the method may further include determining, based on the selection, whether cancer is present in the individual.
[0298] These embodiments for determining the ploidy of chromosome segments may further include detecting single nucleotide variants at single nucleotide variant locations within a set of single nucleotide variant locations, where the detection of chromosomal aneuploidy and / or the single nucleotide variant indicates the presence of circulating tumor nucleic acids in the sample.
[0299] These embodiments may further include receiving haplotype information of the chromosomal segments relating to the tumor of the individual, and using the haplotype information to create a set of models of different ploidy states and allele mismatch rates relating to the set of polymorphic loci.
[0300] As disclosed herein, certain embodiments of a method for determining ploidy may further include excluding outliers from the initial or modified allele frequency data before comparing the initial or modified allele frequencies with the set of models. For example, in certain embodiments, locus allele frequencies that are at least two or three standard deviations greater or less than the mean values of other loci in the chromosome segment are excluded from the data before being used for modeling.
[0301] As is evident from the description herein, many of the embodiments provided herein involve methods for determining the ploidy of chromosome segments, preferably using incompletely phase-determined or fully phase-determined data. Naturally, the methods provided herein also have many features that are improvements over conventional methods for detecting ploidy, and many different combinations of these features can be used.
[0302] In certain aspects of this specification, computer systems and computer-readable media are provided for performing any method of the present invention, as shown in Figures 69 to 70. These include systems and computer-readable media for performing methods for determining ploidy. Thus, in aspects that are non-limiting examples of systems, it is shown that any of the methods provided herein can be performed using systems and computer-readable media that utilize the present disclosure. In other aspects of this specification, systems for detecting chromosomal ploidy in individual samples are provided. Such systems include: a. An input processor configured to receive allele frequency data including the amount of each allele present at each locus in the set of polymorphic loci in the chromosome segment in the sample, b. Modeler, i. By estimating the phase of the allele frequency data, phase-determined allele information relating to the set of polymorphic loci is created. ii. Using the allele frequency data, create individual probabilities of the allele frequencies of the polymorphic loci for different ploidy states. iii. A modeler configured to create a joint probability relating to the set of polymorphic loci using the individual probabilities and the phase-determined allele information, c. Includes a hypothesis management unit configured to determine the ploidy of the chromosome segment by selecting the optimal model that exhibits chromosome ploidy based on the aforementioned joint probability.
[0303] In a particular embodiment of this system, the allele frequency data is data generated by a nucleic acid sequencing system. In a particular embodiment, the system further includes an error correction unit configured to correct errors in the allele frequency data, and the corrected allele frequency data is used by the modeler to create individual probabilities. In a particular embodiment, the error correction unit corrects bias in allele amplification efficiency. In a particular embodiment, the modeler creates the individual probabilities using a set of models for both different ploidy states and allele imbalance rates relating to the set of polymorphic loci. In a particular exemplary embodiment, the modeler creates the joint probabilities by considering linkage between polymorphic loci in the chromosome segment.
[0304] In one exemplary aspect of this specification, a system for detecting chromosome ploidy in an individual sample is provided, which includes the following: a. An input processor configured to receive nucleic acid sequence data relating to alleles present in a set of polymorphic loci in a chromosomal segment of the individual, and to detect the allele frequency in the set of polymorphic loci using the nucleic acid sequence data, b. An error correction unit configured to correct errors in the detected allele frequencies and create corrected allele frequencies for the set of polymorphic loci, c. Modeler, i. By estimating the phase of the nucleic acid sequence data, phase-determined allele information relating to the set of polymorphic loci is created. ii. By comparing the phase-determined allele information with a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci, individual probabilities of allele frequencies for the polymorphic loci for different ploidy states are created. iii. A modeler configured to create a joint probability relating to the set of polymorphic loci by combining the individual probabilities while considering the relative distances between polymorphic loci in the chromosome segment, d. A hypothesis management unit configured to select the optimal model showing chromosomal aneuploidy based on the aforementioned joint probability.
[0305] In certain exemplary systems provided herein, the set of polymorphic loci includes 1,000 to 50,000 polymorphic loci. In certain exemplary systems provided herein, the set of polymorphic loci includes 100 loci known as heterojunction-prone loci. In certain exemplary systems provided herein, the set of polymorphic loci includes 100 loci within 0.5 kb of recombination-prone loci.
[0306] In certain exemplary systems provided herein, the optimal model analyzes the following ploidy states of the first homolog of the chromosome segment and the second homolog of the chromosome segment.
[0307] (1) In which no cell has a deletion or amplification of the first homolog or the second homolog of the chromosome segment,
[0308] (2) A condition in which some or all cells have a deletion of the first homolog of the chromosome segment or an amplification of the second homolog,
[0309] (3) A condition in which some or all cells have a deletion of the second homolog of the chromosome segment or an amplification of the first homolog.
[0310] In certain exemplary systems provided herein, the corrected errors include bias in allele amplification efficiency, impurities, and / or sequencing errors. In certain exemplary systems provided herein, the impurities include environmental impurities and impurities from genotyping analysis. In certain exemplary systems provided herein, the environmental impurities and impurities from genotyping analysis are determined for homozygous alleles.
[0311] In certain exemplary systems provided herein, the hypothesis management unit is configured to analyze the magnitude of the difference between the phase-determined allele information created with respect to the model and the estimated allele frequencies. In certain exemplary systems provided herein, the modeler creates individual probabilities of allele frequencies based on a beta-binomial model of expected and observed allele frequencies in the set of polymorphic loci. In certain exemplary systems provided herein, the modeler creates individual probabilities using a Bayesian classifier.
[0312] In certain exemplary systems provided herein, the nucleic acid sequence data is prepared by performing high-process DNA sequencing on multiple copies of a series of amplification products generated by a multiple amplification reaction, where each of the amplification products is amplified over at least one polymorphic locus of the set of polymorphic loci. In certain exemplary systems provided herein, the multiple amplification reaction is performed under restrictive primer conditions for at least half of these reactions. In certain exemplary systems provided herein, the sample has an average allele mismatch rate of 0.4% to 5%.
[0313] In a particular exemplary system provided herein, the sample is a plasma sample derived from an individual suspected of having cancer, and the hypothesis control unit is further configured to determine, based on the optimal model, whether copy number diversity exists in the tumor cells of the individual.
[0314] In certain exemplary systems provided herein, the sample is a plasma sample derived from an individual, and the hypothesis control unit is further configured to determine, based on the optimal model, whether the individual has cancer. In these embodiments, the hypothesis control unit may further be configured to detect a single nucleotide variant located at a single nucleotide variant position in a set of single nucleotide variant positions, and the detection of chromosomal aneuploidy, a single nucleotide variant, or both indicates the presence of circulating tumor nucleic acid in the sample.
[0315] In certain exemplary systems provided herein, the input processor is further configured to receive haplotype information of the chromosomal segments relating to the tumor of the individual, and the modeler is configured to use the haplotype information to create a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci.
[0316] In certain exemplary systems provided herein, the modeler creates models over allele imbalance rates ranging from 0% to 25%.
[0317] It is evident that any of the methods provided herein can be performed by computer-readable code stored on a non-transient computer-readable medium. Accordingly, one embodiment provided herein is a non-transient computer-readable medium for detecting chromosome ploidy in a sample of an individual, which, when executed by a processing device, performs the following: a. Allele frequency data is received, which includes the amount of each allele present in the sample at each locus among the set of polymorphic loci in the chromosome segment. b. By estimating the phase of the allele frequency data, we create phase-determined gene information relating to the set of polymorphic loci. c. Using the allele frequency data, create individual probabilities of the allele frequencies of the polymorphic loci for different ploidy states. d. Using the individual probabilities and the phase-determined allele information, create a joint probability for the set of polymorphic loci. e. The ploidy of the chromosome segment is determined by selecting the optimal model that exhibits chromosomal ploidy based on the joint probability described above.
[0318] In certain embodiments of the computer-readable medium, the allele frequency data is constructed from nucleic acid sequence data. Further embodiments of the computer-readable medium include correcting errors in the allele frequency data and using the corrected allele frequency data for each individual probability generation step. In certain embodiments of the computer-readable medium, the corrected error is a bias in allele amplification efficiency. In certain embodiments of the computer-readable medium, the individual probabilities are constructed using a set of models for both different ploidy states and allele imbalance rates relating to the set of polymorphic loci. In certain embodiments of the computer-readable medium, the joint probabilities are constructed by considering linkage between polymorphic loci in the chromosome segments.
[0319] A particular embodiment provided herein is a non-transient, computer-readable medium for detecting chromosome ploidy in a sample of an individual, which, when executed by a processing device, performs the following: a. Receiving nucleic acid sequence data relating to alleles present in a set of polymorphic loci in the chromosomal segments of the aforementioned individual, b. Using the nucleic acid sequence data, detect the allele frequency in the set of polymorphic loci, c. Correct the bias in the allele amplification efficiency of the detected allele frequencies to create corrected allele frequencies for the set of polymorphic loci, d. By estimating the phase of the nucleic acid sequence data, phase-determined allele information relating to the set of polymorphic loci is created. e. By comparing the modified allele frequencies with a set of models of different ploidy states and allele imbalance rates for the set of polymorphic loci, we can create individual probabilities of allele frequencies for the polymorphic loci for different ploidy states. f. By combining the individual probabilities while considering the linkage between polymorphic loci in the chromosome segment, a joint probability relating to the set of polymorphic loci is created. g. Based on the above joint probability, select the optimal combination model that exhibits chromosomal aneuploidy.
[0320] In a particular exemplary embodiment of a computer-readable medium, the selection is performed by analyzing the magnitude of the difference between the phase-determined allele information created with respect to the model and the estimated allele frequency.
[0321] In certain exemplary embodiments of computer-readable media, the individual probabilities of the allele frequencies are determined based on a beta-binomial model of expected and observed allele frequencies in the set of polymorphic loci.
[0322] Furthermore, any of the methods of the embodiments provided herein can be performed by executing code stored on a non-transient computer-readable medium.
[0323] Examples of embodiments for detecting cancer In certain aspects, the present invention provides a method for detecting cancer. The sample may be a tumor sample or a liquid sample (e.g., plasma) derived from an individual suspected of having cancer. The method is particularly effective for detecting gene mutations such as single nucleotide variants (e.g., SNVs) or copy number variants (e.g., CNVs) in samples that have small amounts of these gene changes as fractions of total DNA in the sample. Therefore, it has exceptionally high sensitivity for detecting cancer-derived DNA or RNA in the sample. The method can achieve this exceptionally high sensitivity by combining any or all of the improvements for detecting CNVs and SNVs provided herein.
[0324] Accordingly, in certain aspects of this specification, a method for determining whether circulating tumor nucleic acids are present in a sample of an individual, and a non-transient computer-readable medium containing computer-readable code that, when executed by the processing apparatus, causes the processing apparatus to perform the method, are provided. The method comprises the following steps: c. Analyze the sample to determine the ploidy of the set of polymorphic loci in the chromosome segments of the individual, d. Based on the determination of ploidy, the mean allele-misequence rate present at the polymorphic locus is determined, and mean allele-misequence rates of 0.4%, 0.45%, 0.5%, 0.6%, 0.7%, 0.75%, 0.8%, 0.9%, or 1% or more indicate the presence of circulating tumor nucleic acids (e.g., ctDNA) in the sample.
[0325] In some cases, an average allele disequilibrium rate greater than 0.4, 0.45, or 0.5% indicates the presence of ctDNA. In certain embodiments, a method for determining the presence of circulating tumor nucleic acids further comprises detecting a single nucleotide variant at a certain single nucleotide variant position within a set of single nucleotide variant positions, and detecting an allele disequilibrium rate greater than 0.5%, and either or both of the detection of the single nucleotide variant, indicates the presence of circulating tumor nucleic acids in the sample. Naturally, the allele disequilibrium rate, usually the average allele disequilibrium rate, can be determined using any of the methods provided herein for detecting chromosomal ploidy or CNVs. Furthermore, a single nucleotide relating to this aspect of the present invention can be detected using any of the methods provided herein for detecting SNVs.
[0326] In certain embodiments, the method for determining the presence of the circulating tumor nucleic acid further comprises performing the method on a control sample having a known mean allele disequilibrium ratio. For example, the control may be a sample derived from the tumor of the individual. In some embodiments, the control has a mean allele disequilibrium ratio expected for the sample under analysis. For example, the AAI is 0.5% to 5%, or the mean allele disequilibrium ratio is 0.5%.
[0327] In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing 1,000 to 50,000 polymorphic loci for ploidy, or 100 to 1,000 polymorphic loci. In certain embodiments, the analytical step of a method for determining the presence of circulating tumor nucleic acids includes analyzing 100 to 1,000 single nucleotide diversity sites. For example, in these embodiments, the analytical step may include performing multiplex PCR to amplify amplification products across 1,000 to 50,000 synonymous loci and 100 to 1,000 single nucleotide diversity sites. This multiplex reaction can be set up as a single reaction or as a pool of various multiplex sub-reactions. The multiplex reaction methods provided herein (e.g., large multiplex PCR disclosed herein) provide exemplary processes for performing amplification reactions to aid in multiplexing and thus improve sensitivity.
[0328] In certain embodiments, in at least 10%, 20%, 25%, 50%, 75%, 90%, 95%, 98%, 99%, or 100% of the multiplex PCR reactions, the reaction is carried out under restrictive primer conditions. Improved conditions for carrying out large multiplex reactions provided herein may be used.
[0329] In certain aspects, methods for determining whether the circulating tumor nucleic acids are present in a sample of an individual, and all embodiments thereof, can be carried out using a system. This disclosure provides teachings relating to specific functional and structural features for carrying out the methods. In non-limiting examples, the system includes: a. An input processor that analyzes data derived from the sample and determines the ploidy of the set of polymorphic loci in the chromosome segments of the individual, b. A modeler for determining the allele unequilibrium rate present at the polymorphic locus based on the determination of ploidy, wherein an allele unequilibrium rate of 0.5% or more indicates the presence of circulation. Example of an embodiment for detecting a single nucleotide variety
[0330] In certain aspects of this specification, methods for detecting single nucleotide variants in a sample are provided. The improved methods provided herein can detect SNVs in a sample at detection limits of 0.015, 0.017, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, or 0.5 percent. All embodiments of SNV detection can be performed using a system. This disclosure provides teachings relating to specific functional and structural features for performing the methods. Furthermore, embodiments are provided herein that include a non-transient computer-readable medium containing computer-readable code, which, when executed by a processing device, enables the processing device to perform the SNV detection methods provided herein.
[0331] Therefore, one embodiment provided herein is a method for determining whether a single nucleotide variant exists in a set of genomic locations contained in a sample of an individual, the method being: a. For each genome location, use the training dataset to create estimates of the efficiency and error rate per cycle for amplification products across that genome location. b. Obtaining observed nucleotide identity information regarding the location of each genome contained in the sample, c. Using the aforementioned estimates of amplification efficiency and error rate per cycle for each genome location, and comparing the observed nucleotide identity information for each genome location with models of different mutation rates, a set of probabilities of single nucleotide variant rates resulting from one or more actual mutations at each genome location is determined. d. This includes determining the most likely actual mutation rate and confidence level from the set of probabilities for each genome location.
[0332] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the estimates of efficiency and error rate per cycle are made with respect to a series of amplification products across the genomic location. For example, these may include 2, 3, 4, 5, 10, 15, 20, 25, 50, 100, or more amplification products across the genomic location.
[0333] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the observed nucleotide identity information includes the observed total number of reads for each genomic location and the observed number of reads for each diversity allele for each genomic location.
[0334] In an exemplary embodiment of the method for determining the presence of the single nucleotide variant, the sample is a plasma sample and the single nucleotide variant is present in the circulating tumor DNA of the sample.
[0335] In other aspects of this specification, a method is provided for estimating the percentage of single nucleotide manifolds present in a sample derived from a given individual. The method comprises the following steps: a. For a set of genome locations, use the training dataset to create estimates of the efficiency and error rate per cycle for one or more amplification products across those genome locations. b. Obtaining observed nucleotide identity information regarding the location of each genome contained in the sample, c. Using the amplification efficiency of the amplification product and the estimated error rate per cycle, estimates of the mean and variance of the total number of molecules in a survey area with a certain initial actual mutation rate, background error molecules, and actual mutation molecules are created. d. Determine the most likely actual single nucleotide variant rate of the sample resulting from actual mutations by fitting the distribution using the estimated mean and variance of the sample to the observed nucleotide identity information of the sample.
[0336] One example of a method for estimating the percentage of single nucleotide variants present in a sample derived from a particular individual is that the sample is a plasma sample, and the single nucleotide variants are present in the circulating tumor DNA of the sample.
[0337] The training dataset in this embodiment of the present invention typically includes samples from healthy individuals, preferably from a healthy population. In one exemplary embodiment, the training dataset is analyzed on the same day as, or in the same operation as, the analysis of one or more test samples. For example, a training dataset can be created using samples from 2, 3, 4, 5, 10, 15, 20, 25, 30, 36, 48, 96, 100, 192, 200, 250, 500, 1000, or more healthy individuals. When a large number of healthy individuals (e.g., 96 or more) are available in the data, the reliability of the amplification efficiency estimate is increased even if the test samples are tested before performing the method. Since the PCR error rate is per amplification product, nucleic acid sequence data created for the entire amplification region around the SNV can be used, as well as nucleic acid sequence information created for the base position of the SNV. For example, using samples from 50 individuals and sequencing of the amplification product of 20 base pairs around the SNV, the error frequency rate can be determined using error frequency data derived from 1000 base reads.
[0338] Typically, amplification efficiency is estimated by estimating the mean and standard deviation of the amplification efficiency of the amplified segments and fitting them to a distribution model (e.g., a binomial distribution or a beta-binomial distribution). The error rate is determined for a PCR reaction with a known number of cycles, and the error rate per cycle is estimated.
[0339] In certain exemplary embodiments, estimating the starting molecule of the test data set further includes updating the efficiency estimate of the test data set using the starting molecule estimated in step (b) if the number of read observations differs significantly from the estimated number of reads. The estimate may be updated for each new efficiency and / or starting molecule.
[0340] The search space used to estimate the total number of molecules, background error molecules, and actual mutant molecules may include a search space where 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2.5%, 5%, 10%, 15%, 20%, or 25% (lower limit) to 1%, 2%, 2.5%, 5%, 10%, 12.5%, 15%, 20%, 25%, 50%, 75%, 90%, or 95% (upper limit) of the base copies at a given SNV position are SNV bases. In one example where the method detects circulating tumor DNA, lower ranges of 0.1%, 0.2%, 0.25%, 0.5%, and 1% (lower limit) and 1%, 2%, 2.5%, 5%, 10%, 12.5%, or 15% (upper limit) can be used for plasma samples. Higher ranges are used for tumor samples.
[0341] The distribution fits the total number of error molecules (background errors and actual mutations) in the total numerator to calculate the likelihood, or probability, of each possible actual mutation in the search space. This distribution may be a binomial distribution or a beta-binomial distribution.
[0342] The most likely actual mutation is determined by determining the percentage of the most likely actual mutation and calculating the confidence level using data derived from fitting the distribution. As a specific example not intended to limit the clinical interpretation of the method provided herein, a high median mutation rate results in a lower confidence level (percentage) required to determine an SNV as positive. For example, if the median mutation rate of SNVs in a sample using the most likely hypothesis is 5% and the confidence level (percentage) is 99%, a positive SNV call may be made. On the other hand, in this example, if the median mutation rate of SNVs in a sample using the most likely hypothesis is 1% and the confidence level (percentage) is 50%, a positive SNV call may not be made in certain circumstances. Naturally, the clinical interpretation of the data depends on sensitivity, specificity, prevalence, and the availability of alternatives.
[0343] In one exemplary embodiment, the sample is a circulating DNA sample (for example, a circulating tumor DNA sample).
[0344] Other embodiments of this specification provide a method for detecting one or more single nucleotide manifolds present in a test sample of an individual. The method according to this embodiment includes the following steps: d. Based on the results generated during sequencing, the median diversity allele frequency of multiple control samples obtained from each of multiple normal individuals is determined for each single nucleotide diversity position included in the set of single nucleotide diversity positions. This identifies selected single nucleotide diversity positions that have a median diversity allele frequency of normal samples that is below the threshold. After excluding abnormal samples for each of the single nucleotide diversity positions, the background error for each of the single nucleotide diversity positions is determined. e. Based on the data generated during sequencing of the test sample, the weighted mean and variance based on the observed read depth are determined for the selected single nucleotide diversity position of the test sample. f. Detecting one or more single nucleotide variants by using a computer to identify one or more single nucleotide variant locations in which the weighted average based on read depth is statistically significant compared to the background error.
[0345] In a particular embodiment, in the method for detecting one or more SNVs, the sample is a plasma sample, the control sample is a plasma sample, and the detected one or more single nucleotide variants are present in the circulating tumor DNA of the sample. In a particular embodiment, in the method for detecting one or more SNVs, the plurality of control samples comprises at least 25 samples. In one exemplary embodiment, the plurality of control samples has a lower limit of at least 5, 10, 15, 20, 25, 50, 75, 100, 200, or 250 samples, and an upper limit of 10, 15, 20, 25, 50, 75, 100, 200, 250, 500, and 1000 samples.
[0346] In a particular embodiment, in the method for detecting one or more SNVs, the weighted average and observed variance by the observed read depth are calculated by excluding outliers from the data created during high-process sequencing. In a particular embodiment, in the method for detecting one or more SNVs, the read depth for each single nucleotide diversity position in the test sample is at least 100 reads.
[0347] In certain embodiments, the method for detecting one or more SNVs includes a multiple amplification reaction performed under restrictive primer reaction conditions, wherein the sequencing is performed. These exemplary embodiments are performed using an improved method for performing the multiple amplification reaction provided herein.
[0348] While not limited by theory, the method of this embodiment utilizes a background error model using a normal plasma sample sequenced in the same sequencing operation as the test sample to account for specific artificial products in practice. Locations suspected to be noise with normal central diversity allele frequencies greater than thresholds (e.g., greater than 0.1%, 0.2%, 0.25%, 0.5%, 0.75%, and 1.0%) are excluded.
[0349] To account for noise and impurities, samples with outlier values are repeatedly excluded from the model. For each base substitution at each gene locus, the weighted mean and standard deviation of the errors are calculated based on the read depth. In certain exemplary embodiments, samples (e.g., tumor or cell-free plasma samples) with at least a threshold number of reads (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1000 variant reads) and single nucleotide diversity locations with Z-values greater than 2.5, 5, 7.5, or 10 against the background error model in a particular embodiment are measured as candidate mutations.
[0350] In certain embodiments, in the sequencing of each single nucleotide diversity position, read depths are achieved such that the lower limit is 100, 250, 500, 1,000, 2,000, 2,500, 5,000, 10,000, 20,000, 25,0000, 50,000, or more than 100,000 reads, and the upper limit is 2,000, 2,500, 5,000, 7,500, 10,000, 25,000, 50,000, 100,000, 250,000, or more than 500,000 reads. The sequencing operation is typically a high-processing sequencing operation. In exemplary embodiments, the mean or median values generated for the test sample are weighted by read depth. Therefore, the likelihood that a particular diversity allele will be determined in a sample where one diversity allele is detected per 1,000 reads is weighted higher than the likelihood that one diversity allele will be detected per 10,000 reads. Since the determination of a particular diversity allele (i.e., a mutation) cannot be performed with 100% confidence, the identified single nucleotide variant can be considered a candidate variant or a candidate mutation.
[0351] Examples of test statistics for analyzing phase-determined data The following are examples of test statistics for the analysis of phase-determined data from a mixed sample that is known or suspected to contain DNA or RNA from two or more genetically non-identical cells. Let f be the fraction of target DNA or RNA (e.g., the DNA or RNA fraction with the target CNV, or the DNA or RNA fraction derived from the target cell, such as a cancer cell). In some embodiments, in the case of prenatal testing, f refers to the fraction of fetal DNA, RNA, or cells in a mixture of fetal and maternal DNA, RNA, or cells. Assuming that two copies of DNA were produced by each of the target cells, f refers to the fraction of DNA derived from the target cell. f is different from the fraction of DNA derived from the target cell that exists in a deleted or duplicated segment.
[0352] Let A and B be the possible allele values for each SNP. All possible ordered allele pairs are shown using AA, AB, BA, and BB. In some embodiments, SNPs with ordered alleles AB or BA are analyzed. i Let A be the number of sequence reads for the i-th SNP. i and B i Let these be the number of reads for the i-th SNP representing alleles A and B, respectively. Assume the following: N i =A i +B i Allergen ratio R i It is defined as follows:
number
[0353] Without loss of generality, some embodiments focus on a single chromosome segment. For further clarity, the term “first homologous chromosome segment compared to second homologous chromosome segment” means the first homolog of a given chromosome segment and its second homolog. In some such embodiments, all of the SNPs of interest are contained within the segment chromosome of interest. In other embodiments, multiple chromosome segments are analyzed for possible copy number variations.
[0354] Map Estimation This method uses knowledge of phase determination via ordered alleles to detect deletions or duplications in the target segment. Each SNPi is defined as follows:
number
[0355] Therefore, it is defined as follows:
number
[0356] Dichromosomal hypothesis In this hypothesis that the target segment has neither deletion nor duplication,
Number
Number
Number
[0357] Deletion Hypothesis In the hypothesis that the first homolog is deleted (i.e., the SNP of AB becomes B and the SNP of BA becomes A), R i has a binomial distribution with parameters {1 - (1 / (2 - f))} and T for the SNP of AB, and (1 / (2 - f)) and T for the SNP of BA. Therefore,
Number
Number
[0358] In the hypothesis that the second homolog is deleted (i.e., the SNP of AB becomes A and the SNP of BA becomes B), R iIt has a binomial distribution with parameters (1 / (2-f)) and T for SNPs in AB, and {1-(1 / (2-f))} and T for SNPs in BA. Therefore,
number
number
[0359] Overlap Hypothesis The hypothesis that the first homolog is duplicated (i.e., the SNP of AB becomes AAB and the SNP of BA becomes BBA) is that R i It has a binomial distribution with parameters (1+f) / (2+f) and T for SNPs in AB, and {1-((1+f) / (2+f))} and T for SNPs in BA. Therefore,
number
number
[0360] The hypothesis that the second homolog is duplicated (i.e., the SNP of AB becomes ABB and the SNP of BA becomes BAA) is that R i It has a binomial distribution with parameters {1-((1+f) / (2+f))} and T for SNPs in AB, and (1+f) / (2+f) and T for SNPs in BA. Therefore,
number
number
[0361] classification As shown in the section above, X i This is the following binomial random variable:
number
[0362] This allows us to calculate the probability of the test statistic S for each hypothesis. We can calculate the probability of each hypothesis assuming the measured data. In some embodiments, the hypothesis with the highest probability is selected. If necessary, each N is defined by the read depth constant N. i The distribution of S can be simplified by approximating it or by truncating the read depth to a constant N. This simplification yields the following:
number
[0363] The value of f can be estimated by selecting the most likely value of f, such as the value that best fits the data, using algorithms (e.g., search algorithms) such as maximum likelihood estimation, maximum posterior probability estimation, or Bayesian estimation, assuming the measured data. In some embodiments, multiple chromosomal segments are analyzed and the value of f is estimated based on the data of each segment. If all of the target cells have these overlaps or deletions, the estimates of f based on the data of these different segments will be similar. In some embodiments, f is measured experimentally by determining the DNA or RNA fraction derived from cancer cells based on the difference in methylation (hypomethylation or hypermethylation) between cancer and non-cancerous DNA or RNA.
[0364] In some embodiments, in the case of a mixed sample of fetal nucleic acids and maternal nucleic acids, the value of f is the fetal fraction, i.e., the fraction of fetal DNA (or RNA) relative to the total amount of DNA (or RNA) in the sample. In some embodiments, the fetal fraction is determined by obtaining genotype data from a maternal blood sample (or fraction thereof) for a set of polymorphic loci on at least one chromosome where dichromacy is predicted in both the mother and the fetus; creating multiple hypotheses corresponding to different possible fetal fractions for the chromosome; constructing a model for expected allele measurements in the blood sample for the set of polymorphic loci on the chromosome for the possible fetal fractions; calculating the relative probability of each hypothesis regarding the fetal fraction using the model and allele measurements derived from the blood sample or fraction thereof; and selecting the fetal fraction corresponding to the hypothesis with the highest probability to determine the fetal fraction in the blood sample. In some embodiments, the fetal fraction is determined by identifying polymorphic loci in the mother in which the first allele of a certain polymorphic locus is homozygous and in the father in which (i) the first and second alleles of that polymorphic locus are heterozygous, or (ii) the second allele is homozygous, and determining the fetal fraction in the blood sample using the amount of the second allele detected in the blood sample for each identified polymorphic locus (see, for example, U.S. Patent Publication 2012 / 0185176, provided on March 29, 2012, and U.S. Patent Publication 2014 / 0065621, filed on March 13, 2013, the entire contents of which are incorporated herein by reference, respectively).
[0365] Other methods for determining fetal fractions include modeling possible fetal fractions by measuring alleles present at numerous polymorphic (e.g., SNP) gene loci using high-speed DNA sequencing equipment (see, for example, U.S. Patent Publication 2012 / 0264121, which is incorporated herein by reference in its entirety). Other methods for calculating fetal fractions are also found in Sparks et al., "Noninvasive prenatal detection and selective analysis of cell-free DNA obtained from maternal blood: evaluation for trisomy 21 and trisomy 18" Am J Obstet Gynecol 2012;206:319.e1-9, which is incorporated herein by reference in its entirety. In some embodiments, fetal fractions are determined using methylation assays that assume certain loci are methylated, or preferentially methylated, in the fetus, while these same loci are unmethylated, or preferentially unmethylated, in the mother (see, for example, U.S. Patents 7,754,428, 7,901,884, and 8,166,382, the entire contents of which are incorporated herein by reference).
[0366] Figures 1A to 13D are graphs showing the distribution of the test statistic S divided by T (number of SNPs) ("S / T") for different copy number hypotheses for various read depths and tumor fractions as the number of SNPs increases (where f is the tumor DNA fraction in total DNA).
[0367] Rejection of a single hypothesis The distribution of S in the disomal hypothesis is independent of f. Therefore, the probability of the measured data can be calculated for the disomal hypothesis without calculating f. A single hypothesis rejection test can be used for the null hypothesis of disomal hypothesis. In some embodiments, the probability of S in the disomal hypothesis is calculated, and if that probability is below a predetermined threshold (e.g., less than 1 in 1000), the disomal hypothesis is rejected. This indicates the existence of duplication or deletion of the chromosomal segment. The threshold can be adjusted as needed to change the false positive rate.
[0368] Examples of phase-determined data analysis methods The following describes exemplary methods for analyzing data from a sample that is known or suspected to be a mixed sample containing DNA or RNA from two or more genetically non-identical cells. In some embodiments, phase-determined data are used. In some embodiments, the method includes determining whether each calculated allele ratio is greater than or less than the expected allele ratio, and determining the magnitude of the difference at a particular locus. In some embodiments, a likelihood distribution is determined for the allele ratios present at a locus with respect to a particular hypothesis, and the closer the calculated allele ratio is to the center of the likelihood distribution, the more likely the hypothesis is to be correct. In some embodiments, the method includes determining the likelihood that the hypothesis is correct at each locus. In some embodiments, the method includes determining the likelihood that the hypothesis is correct at each locus and combining the probability of the hypothesis at each locus with the hypothesis, and selecting the highest composite probability. In some embodiments, the method includes determining the likelihood that the hypothesis is correct for each locus and for each possible ratio of DNA or RNA from one or more target cells to the total DNA or RNA in the sample. In some embodiments, the combined probability of each hypothesis is determined by combining the probability of each hypothesis at each seat with each possible ratio, and the hypothesis with the highest combined probability is selected.
[0369] In one embodiment, the following hypothesis can be considered: H 11 (All cells are normal), H 10(There are cells that have only homolog 1, and cells that lack homolog 2), H 01 (There are cells of homolog 1 that have only homolog 2), H 21 (Cells with a duplication of homolog 1 exist), H 12 (Cells with homology 2 duplication are present). For a fraction f of target cells such as cancer cells or mosaic cells (or DNA or RNA fraction derived from the target cells), the expected allele ratio of heterozygous (AB or BA) SNPs is obtained as follows. Formula (1)
number
[0370] Correction of bias, impurities, and sequencing errors. Observation D at the aforementioned SNP s Each allele is found in the original mapping read n A 0 and n B 0 This consists of the following. Therefore, the corrected read n using the prediction bias in the amplification of alleles A and B. A and n B It is possible to find this.
[0371] c a Let r(c) be an environmental impurity (for example, an impurity derived from DNA in the air or environment), and r(c a ) is defined as the allergen ratio to environmental impurities (here, it is initially set to 0.5). Also, c g Let r(c) be the impurity rate determined by genotype analysis (e.g., impurities originating from other samples), and r(c) g ) is defined as the allele ratio of impurities. e (A,B) and s e (B,A) represents a sequence determination error in which one allele is mistakenly identified as another allele (for example, by incorrectly detecting allele A when allele B is present).
[0372] After correcting for environmental impurities, impurities from genotyping analysis, and sequencing errors, the observed allele ratio q(r,c) is obtained relative to a given expected allele ratio r. a ,r(ca ), c g ,r(c g ),s e (A,B),s e (B,A) can be found.
[0373] Since the genotype of the impurities is unknown, we will use population frequencies to determine P(r(c g We can then calculate P(r(c g )=0)=(1-p) 2 , P(r(c g )=0)=2p(1-p), and P(r(cg)=0)=p 2 r(c g Using the conditional expectation value for ), E[q(r,c a ,r(c a ), c g ,r(c g ),s e (A,B),s e (B,A)) can be determined. Note that environmental impurities and impurities determined by genotyping are determined using homozygous SNPs and are therefore not affected by the presence or absence of deletions or duplications. Furthermore, if necessary, environmental impurities and impurities determined by genotyping can be measured using a reference chromosome.
[0374] Possibility of each SNP Assuming an allele ratio r, n A and n B The probability of observing this can be calculated. Formula (2):
number
[0375] D s Let the data be SNPs. Each hypothesis h∈{H 11 ,H 01 ,H 10 ,H 21 ,H 12Regarding}, assuming that in equation (1) r=r(AB,h) or r=r(BA,h), r(c g ) calculate the conditional expected value for the observed allele ratio E[q(r,c a ,r(c a ), c g ,r(c g ))] can be determined. Therefore, in equation (2), r = E[q(r,c a ,r(c a ), c g ,r(c g ),s e (A,B),s e Assuming (B,A)), P(D s We can find |h,f)
[0376] Search algorithm In some embodiments, SNPs with allele ratios that appear to be outliers are ignored (for example, by ignoring or excluding SNPs with allele ratios that are at least two or three standard deviations larger or smaller than the mean). The advantage of this method is that, since the variability of the allele ratio may be high when the percentage of mosaicism present is high, the deletion of this SNP due to mosaicism is reliably prevented.
[0377] F={f1,…,f N} is used as the search space for mosaic percentages (e.g., tumor fractions). P(D) of each SNP s By finding |h,f) and f∈F, we can combine the likelihoods of all SNPs.
[0378] The algorithm examines each f for each hypothesis. Using a certain search method, if there is a range F* of f (where the confidence level of the hypothesis with deletions or overlaps is higher than the confidence level of the hypothesis without deletions or overlaps), it is concluded that a mosaic exists. In some embodiments, P(D) at F* s The maximum likelihood estimate of |h,f) is determined. If necessary, the conditional expectation value for f∈F* may also be determined. If necessary, the confidence level of each hypothesis may also be determined.
[0379] Further aspects In some embodiments, a beta-binomial distribution is used instead of a binomial distribution. In some embodiments, specific parameters of the beta-binomial distribution of the sample are determined using a reference chromosome or chromosome segment.
[0380] Theoretical performance using simulations If necessary, the theoretical performance of the algorithm may be evaluated by randomly assigning a baseline number of reads to SNPs at a given read depth (DOR). Under normal circumstances, the binomial probability parameter p=0.5 is used, and p is revised accordingly for deletions or duplicates. Examples of input parameters for each simulation include (1) the number of SNPs(S), (2) the constant DOR(D) for each SNP, (3) p, and (4) the number of experiments.
[0381] First simulation experiment This experiment focused on S∈{500,1000}, D∈{500,1000}, and p∈{0%,1%,2%,3%,4%,5%}. 1,000 simulation experiments were performed for each setting (24,000 experiments with phase and 24,000 experiments without phase). Read count simulations were performed using a binomial distribution (other distributions could be used if necessary). False positive rates (in this case, p=0%) and false negative rates (in this case, p>0%) were determined with and without phase information. False positive rates are shown in Figure 26. Phase information is particularly useful for S=1000 and D=1000. Even for S=500 and D=500, the algorithm yields the highest false positive rate with or without phase from the given test conditions. False negative rates are shown in Figure 27.
[0382] Phase information is particularly useful when the mosaic percentage is low (≤3%). Without phase information, many false negatives were observed at p=1%. This is because the confidence level regarding deletions is H 10 and H 01The decision is made by allocating equal opportunities to each hypothesis, because a small deviation supporting one hypothesis is insufficient to compensate for the low likelihood from other hypotheses. This also applies to overlaps. The aforementioned algorithm appears to be more sensitive to read depth than to the number of SNPs. Results using phase information assume that complete phase information is available for many consecutive heterojunction SNPs. Haplotype information can be obtained by probabilistically combining haplotypes in smaller segments as needed.
[0383] Second simulation experiment In this experiment, focusing on S∈{100,200,300,400,500}, D∈{1000,2000,3000,4000,5000}, and p∈{0%,1%,1.5%,2%,2.5%,3%}, 10,000 random experiments were conducted for each setting. The false positive rate (in this case, p=0%) and false negative rate (in this case, p>0%) were determined with and without phase information. The false negative rate was less than 10% for D≧3000 and N≧200 using haplotype information, but the same performance was obtained for D=5000 and N≧400 (Figures 20A and 20B). The difference in false negative rates was particularly pronounced when the percentage of mosaicism was small (Figures 21A~25B). For example, at p=1%, a false negative rate of less than 20% cannot be obtained without using haplotype data, but at N≧300 and D≧3000, the false negative rate approaches 0%. At p=3%, a false negative rate of 0% is observed with haplotype data, but without haplotype data, N≧300 and D≧3000 are required to obtain the same performance.
[0384] Examples of methods for detecting deletions and duplicates without using phase-determined data. In some embodiments, unphased genetic data is used to determine whether there is copy number overpopulation of a first homologous chromosome segment compared to a second homologous chromosome segment in the individual's genome (e.g., the genome of one or more cells, or cell-free DNA or cell-free RNA). In some embodiments, phased genetic data is used, but phase determination is ignored. In some embodiments, the DNA or RNA sample is a mixed sample of cell-free DNA or cell-free RNA containing cell-free DNA or cell-free RNA from an individual containing two or more genetically distinct cells. In some embodiments, the method uses the magnitude of the difference between the calculated allele ratio and the expected allele ratio for each of the loci.
[0385] In some embodiments, the method includes measuring the amount of each allele present at each locus to obtain genetic data of a set of polymorphic loci in a chromosome or chromosome segment in a DNA or RNA sample derived from one or more cells from the individual. In some embodiments, an allele ratio is calculated for loci that are heterozygous in at least one cell from which the sample originates (e.g., loci that are heterozygous in the fetus and / or loci that are heterozygous in the mother). In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele by the total measurement of all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is obtained by dividing one measurement of the allele (e.g., an allele in the first homologous chromosome segment) by the measurement of one or more other alleles at that locus (e.g., an allele in the second homologous chromosome segment). The calculated allele ratio and the expected allele ratio may be calculated using any of the methods disclosed herein, or using any standard method (for example, any mathematical transformation of the calculated allele ratio or expected allele ratio described herein).
[0386] In some embodiments, the test statistic is calculated based on the magnitude of the difference between the calculated allele ratio and the expected allele ratio for each of the loci. In some embodiments, the test statistic Δ is calculated using the following formula.
number
[0387] For example, when the expected allele ratio is 0.5, δ i It can be defined as follows:
number
[0388] In some embodiments, a set of one or more hypotheses defining the copy number of a chromosome or chromosome segment in one or more genomes of the cell is enumerated. In some embodiments, the most likely hypothesis is selected based on the test statistic to determine the copy number of a chromosome or chromosome segment in one or more genomes of the cell. In some embodiments, a hypothesis is selected if the probability that the test statistic for that hypothesis belongs to the distribution of the test statistic is greater than an upper threshold, and one or more of the hypotheses are rejected if the probability that the test statistic for that hypothesis belongs to the distribution of the test statistic is less than a lower threshold. Alternatively, if the probability that the test statistic for that hypothesis belongs to the distribution of the test statistic is between the lower and upper thresholds, or if the probability has not been determined with sufficient confidence, the hypothesis is neither selected nor rejected. In some embodiments, the upper and / or lower thresholds are determined from an experimental distribution (e.g., a distribution derived from training data (e.g., a sample with a known copy number, such as a diploid sample, or a sample known to have a specific deletion or duplication)). Such an experimental distribution can be used to select the thresholds for a single-hypothesis rejection test.
[0389] Furthermore, the test statistic Δ is independent of S, and therefore, both can be used independently as needed.
[0390] Examples of methods for detecting deletions and duplications using allele distributions or patterns. This section includes a method for determining whether there is copy number overpopulation of a first homologous chromosome segment compared to a second homologous chromosome segment. In some embodiments, the method includes (i) listing several hypotheses specifying the copy number of a chromosome or chromosome segment present in the genome of one or more cells (e.g., cancer cells) of the individual, or (ii) defining several hypotheses defining the degree of copy number overpopulation of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells of the individual. In some embodiments, the method includes obtaining genetic data derived from the individual for several polymorphic loci (e.g., SNP loci) in the chromosome or chromosome segment. In some embodiments, for each of the hypotheses, a probability distribution of the expected genotype of the individual is constructed. In some embodiments, the data fit between the obtained genetic data of the individual and the probability distribution of the expected genotype of the individual is calculated. In some embodiments, one or more hypotheses are ranked according to the data fit, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm (e.g., a search algorithm) is used in one or more of the steps of calculating the data fit, ranking the hypotheses, and selecting the highest-ranked hypothesis. In some embodiments, the data fit is a fit to a beta-binomial distribution or a fit to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, maximum posterior probability estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained gene data and the expected gene data.
[0391] In some embodiments, the method includes (i) listing several hypotheses specifying the copy number of a chromosome or chromosome segment present in the genome of one or more cells (e.g., cancer cells) of the individual, or (ii) defining the degree of over-occurrence of the copy number of a first homologous chromosome segment compared to a second homologous chromosome segment in the genome of one or more cells of the individual. In some embodiments, the method includes obtaining genetic data derived from the individual for several polymorphic loci (e.g., SNP loci) in the chromosome or chromosome segment. In some embodiments, the genetic data includes the number of alleles for the several polymorphic loci. In some embodiments, for each hypothesis, a joint distribution model is constructed for the expected number of alleles in the several polymorphic loci in the chromosome or chromosome segment. In some embodiments, the relative probabilities of one or more hypotheses are determined using the joint distribution model and the number of alleles measured in the sample, and the hypothesis with the highest probability is selected.
[0392] In some embodiments, the presence or absence of CNVs (e.g., deletions or duplications) is determined using the distribution or pattern of alleles (e.g., patterns of calculated allele ratios). If necessary, the parental origin of the CNV may be determined based on this pattern. A maternal duplication is an extra copy of a maternal chromosome segment, while a maternal deletion is a missing maternal copy of a chromosome segment, with the only existing copy of the segment originating from the father. Examples of patterns are shown in Figures 15A–19D and further explained below.
[0393] To determine the presence or absence of a deletion in a target chromosome segment, the algorithm considers the distribution of sequence measurements originating from each of two possible alleles for a number of SNPs per chromosome. Importantly, some aspects of the algorithm employ methods unsuitable for visualization. Therefore, for illustrative purposes, the data is simply shown in Figures 15A to 18 as the ratio of the two most likely alleles A and B, allowing for a clearer visualization of the relevant trends. This simplified example does not consider some of the possible characteristics of the algorithm. For example, two aspects of the algorithm that cannot display the data using any visualization method that shows allele ratios are: 1) the ability to utilize linkage disequilibrium, i.e., the influence of a measurement of one SNP on the likelihood of identity of neighboring SNPs, and 2) the use of a non-Gaussian data model that explains the expected distribution of allele measurements of SNPs assuming platform characteristics and amplification bias. Also, a simplified version of the algorithm considers only the most common allele for each SNP, ignoring other possible alleles.
[0394] The target deletion was detected in genomic samples and maternal blood samples. In some embodiments, genomic samples and maternal plasma samples were analyzed using the multiplex PCR and sequencing method of Example 1. The tested genomic DNA syndrome samples did not contain heterozygous SNPs in the target region, confirming that this assay can distinguish between monochromosomal (symptomatic) and dichromosomal (non-symptomatic) syndromes. Analysis of cell-free DNA from maternal blood samples was able to detect 22q11.2 deletion syndrome, cat cry deletion syndrome, Wolf-Hirschhorn deletion syndrome, and other deletion syndromes in the fetus shown in Figure 14.
[0395] Figures 15A-15C show data indicating the presence or absence of two chromosomes in the following cases: when the sample is entirely maternal (no fetal cell-free DNA is present) (Figure 15A), when it contains a moderate concentration (12%) of fetal cell-free DNA fraction (Figure 15B), and when it contains a high concentration (26%) of fetal cell-free DNA fraction (Figure 15C). The x-axis represents the linear position of the individual's polymorphic locus along the chromosome, and the y-axis represents the number of reads for allele A as a fraction of the total number of reads for allele (A+B). The maternal and fetal genotypes are shown to the right of the plot. Each plot is color-coded according to the maternal genotype. For example, red represents maternal genotype AA, blue represents maternal genotype BB, and green represents maternal genotype AB. The measurement was performed on total cell-free DNA isolated from the mother's blood, and this cell-free DNA contains cell-free DNA from both the mother and the fetus. Therefore, each point indicates the combination of fetal and maternal DNA that produced that SNP. Therefore, as the percentage of maternal cell-free DNA increases from 0% to 100%, some points in the plot will fluctuate slightly depending on the maternal and fetal genotypes.
[0396] In all cases, SNPs where allele A is homozygous (AA) in both mother and fetus are closely associated with the upper limit of the plot, because the fraction of allele A reads is high due to the absence of allele B. Conversely, SNPs where allele B is homozygous in both mother and fetus are closely associated with the lower limit of the plot, because the fraction of allele A reads is low due to the presence of only allele B. Points not closely associated with the upper and lower limits of the plot represent SNPs where the mother, fetus, or both are heterozygous. These points are useful for identifying fetal deletions or duplications, but may also be useful for determining parental and maternal inheritance. These points are separated based on both maternal and fetal genotypes and fetal fractions, and the exact position of each point along the y-axis depends on both stoichiometry and fetal fractions. For example, a locus where the mother is AA and the fetus is AB is predicted to have different fractions of allele A reads, and therefore its position on the y-axis differs due to the fetal fraction.
[0397] Figure 15A shows data from non-pregnant women, illustrating a pattern where the genotype is entirely maternal. This pattern includes clusters of points. The red clusters are closely associated with the top of the plot (here, SNPs with maternal genotype AA), the blue clusters with the bottom of the plot (here, SNPs with maternal genotype BB), and the single central green cluster (here, SNPs with maternal genotype AB). In Figure 15B, the positions of some alleles have shifted up and down along the y-axis because the fetal alleles are included in the fraction of allele A reads. In Figure 15C, a pattern is readily apparent, including two red peripheral bands, two blue peripheral bands, and three green bands. These three green bands correspond to SNPs that are heterozygous in the mother, while the two "peripheral" bands at both the top (red) and bottom (blue) of the plot correspond to SNPs that are homozygous in the mother, respectively.
[0398] Figure 16A shows the analysis of a carrier of the 22q11.2 deletion (mother of the carrier). The carrier of this deletion does not have a heterozygous SNP in this region, because only this carrier has one copy of this region. Therefore, this deletion is indicated by the absence of the green A / B SNP. Figure 16B shows the analysis of a 22q11 deletion in a fetus inherited from both parents. This fetus alone inherits a single copy of the chromosome segment (in the case of a deletion inherited from both parents, the copy present in the fetus originates from the mother), and therefore, if a single allele is inherited at each locus of this segment, heterozygosity in this fetus is impossible. Therefore, the only possible SNP identity for this fetus is A or B. Note that the inner part of the peripheral band is missing. In the case of deletions inherited from both parents, the characteristic pattern includes two green bands indicating heterozygous SNPs in the mother, and single red and blue bands indicating homozygous SNPs in the mother, closely associated with the upper and lower limits (1 and 0) of the plot, respectively.
[0399] Figure 17 shows an analysis of cat cry deletion syndrome inherited from the mother. Instead of three green bands, there are two green bands, two red peripheral bands, and two blue peripheral bands. Since both the mother and fetus have the deletion, the deletion inherited from the mother (e.g., a carrier of a mother with Duchenne muscular dystrophy) can also be detected based on a small signal in the deletion region of a mixed maternal and fetal DNA sample (e.g., plasma sample).
[0400] Figure 18 is a plot of Wolf-Hirschhorn deletion syndrome inherited from both parents, indicated by one red peripheral band and one blue peripheral band.
[0401] If necessary, similar plots can be created for samples from individuals suspected of having deletions or duplications (e.g., CNVs associated with cancer). In such plots, the following color codes can be used based on the genotype of cells without CNVs: red indicates the AA genotype, blue indicates the BB genotype, and green indicates the AB genotype. In some embodiments, in the case of deletions, the pattern includes two green bands representing SNPs that are heterozygous in the individual (the upper green band shows AB from cells without deletions and A from cells with deletions, and the lower green band shows AB from cells without deletions and B from cells with deletions), and a single red band and a blue band representing SNPs that are homozygous in the individual and are closely associated with the upper and lower limits (1 and 0) of the plot, respectively. In some embodiments, the separation of the two green bands also increases as the fraction of cells, DNA, or RNA with deletions increases.
[0402] Examples of methods for identifying and analyzing multiple pregnancies In some embodiments, the presence of a multiple pregnancy, such as a twin pregnancy, is detected using any of the methods of the present invention, wherein at least one of the fetuses is genetically distinct from at least one other fetus. In some embodiments, dizygotic twins are identified based on the presence of two fetuses having different alleles, different allele ratios, or different allele distributions at some (or all) of the loci examined. In some embodiments, dizygotic twins are identified by determining the expected allele ratios for each locus (e.g., SNP locus) for two fetuses that may have the same or different fetal fractions in a sample (e.g., a plasma sample). In some embodiments, the likelihood of a particular pair of fetal fractions (where f1 is the fetal fraction of fetus 1 and f2 is the fetal fraction of fetus 2) is calculated, taking into account some or all of the possible genotypes of a particular pair of fetuses, given the maternal genotype and genotype population frequency. The expected allele ratio of a given SNP is determined by combining a mixture of the genotypes of the two fetuses and the genotype of one mother with the fetal fractions. For example, if the mother is AA, fetus 1 is AA, and fetus 2 is AB, the total fraction of allele B in this SNP is half that of f2. In the likelihood calculation, the degree to which all SNPs fit the expected allele ratio is calculated based on all possible combinations of fetal genotypes. The pair of fetal fractions (f1, f2) that best fit the data is selected. It is not necessary to calculate the specific genotype of the fetus. Instead, all possible genotypes that occur in statistical combinations can be considered, for example. In some embodiments, if the above method does not distinguish between mono fetuses and identical twins, ultrasound can be used to determine whether it is a mono fetus or an identical twin pregnancy. If a twin pregnancy is detected by ultrasound, it can be assumed to be an identical twin pregnancy, because if it were a sibling twin pregnancy, it should be detected based on the SNP analysis described above.
[0403] In some embodiments, it is known from pre-pregnancy examinations (e.g., ultrasound) that a pregnant mother is having a multiple pregnancy (e.g., twins). Any of the methods of the present invention can be used to determine whether the multiple pregnancy includes identical twins or sibling twins. For example, the measured allele ratio can be compared with what is expected for identical twins (same allele ratio as a single pregnancy) or fraternal twins (e.g., calculation of the allele ratio as described above). Some identical twins are monochorionic twins and are at risk of twin-to-twin transfusion syndrome. Therefore, it is desirable that twins determined to be identical twins using the methods of the present invention be examined (e.g., by ultrasound) to determine whether they are monochorionic twins, and if they are monochorionic twins, these twins can be monitored (e.g., every other week from 16 weeks by ultrasound).
[0404] In some embodiments, one of the methods of the present invention is used to determine whether one of the fetuses in a multiple pregnancy (e.g., a twin pregnancy) is aneuploid. The twin aneuploidy test begins with an estimate of the fetal fractions. In some embodiments, a pair of fetal fractions (f1, f2) that best fit the data is selected as described above. In some embodiments, maximum likelihood estimation is performed for the parameter pair (f1, f2) over the range of possible fetal fractions. In some embodiments, the range of f2 is from 0 to f1, because f2 is defined as the smaller fetal fraction. Assuming a pair of (f1, f2), the likelihood of the data is calculated from the allele ratios observed in a set of loci, such as SNP loci. In some embodiments, the data likelihood reflects the maternal genotype, the paternal genotype if available, population frequency, and the probability of the resulting fetal genotype. In some embodiments, SNPs are assumed independently. The estimated fetal fraction pair is the one that produces the highest data likelihood. If f2 is 0, this data is best explained by only one set of fetal genotypes indicating identical-sex twins. In this case, f1 is the combined fetal fraction. Otherwise, f1 and f2 are estimates of the fetal fractions for the individual twins. Once the best estimates of (f1, f2) are established, the total fraction of allele B in plasma can be predicted for any combination of maternal and fetal genotypes as needed. It is not necessary to assign individual sequence reads to individual fetuses. Ploidy testing is performed using other maximum likelihood estimates that compare the data likelihoods of the two hypotheses. In some embodiments, for identical-sex twins, the hypotheses (i) that both twins are euploid and (ii) that both twins are trichromosomes are considered. In some embodiments, for dizygotic twins, the hypotheses (i) that both twins are euploid and (ii) that at least one of the twins is trichromosomes are considered. The triploidy hypothesis for dizygotic twins is based on low fetal fractionation, because triploidy can also be detected in twins with high fetal fractionation. Ploidy likelihood is calculated using a method that predicts the expected number of reads at each target genomic locus conditioned on either the diloidy hypothesis or the triploidy hypothesis. There are no requirements for diloidy reference chromosomes.The variance model for the expected number of reads takes into account the performance of individual loci and the correlations between loci (for example, U.S. Serial No. 62 / 008,235 filed on June 5, 2014, and U.S. Serial No. 62 / 032,785 filed on August 4, 2014, the full contents of which are incorporated herein by reference). If small twins have fetal fraction f1, the ability to detect the triploidy of those twins is equivalent to the ability to detect triploidy in a singleton pregnancy with the same fetal fraction. This is because, in some embodiments, some of the methods for detecting triploidy are genotype-independent and do not distinguish between multiple and singleton pregnancies. It simply searches for an increase in the number of reads according to the determined fetal fraction.
[0405] In some embodiments, the method includes detecting the presence of twins based on SNP loci (for example, as described above). If twins are detected, the fetal fractions of each fetus (f1, f2) are determined using the SNPs as described above. In some embodiments, the amplification bias per SNP is determined using samples with high-confidence disomal calling. In some embodiments, these samples with high-confidence disomal calling are analyzed in the same procedure as one or more samples of interest. In some embodiments, the amplification bias per SNP is used to model the read distribution for one or more chromosomes or segments of interest (e.g., chromosome 21) or for disomal and trisomal hypotheses, assuming the lower of the fetal fractions of the two twins. The likelihood, or probability, of disomal or trisomal is calculated based on the two models and the measured values of the chromosome or segment of interest.
[0406] In some embodiments, the threshold for a positive aneuploidy call (e.g., a trisomal call) is set based on the twin having the lower fetal fraction. In this case, if the other twin is positive, or both are positive, the total chromosome occurrence is clearly greater than the threshold.
[0407] Examples of measurement / quantification methods In some embodiments, one or more CNSs (e.g., deletions or duplications of chromosome segments or entire chromosomes) are detected using one or more measurement methods (also called quantitative methods). In some embodiments, one or more measurement methods are used to determine whether the copy number overpopulation of the first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of the second homologous chromosome segment. In some embodiments, one or more measurement methods are used to determine the number of duplicated chromosome segments or chromosomes (e.g., whether there are one, two, three, four, or more extra copies). In some embodiments, one or more measurement methods are used to distinguish between samples with many duplications and the smaller tumor fraction and samples with fewer duplications and the larger tumor fraction. For example, one or more measurement methods may be used to distinguish between a sample with four extra copies of a chromosome and a 10% tumor fraction and a sample with two extra copies of a chromosome and a 20% tumor fraction. Exemplary methods are disclosed, for example, in U.S. Publications 2007 / 0184467, 2013 / 0172211, and 2012 / 0003637, U.S. Patents 8,467,976, 7,888,017, 8,008,018, 8,296,076, and 8,195,415, U.S. Serial No. 62 / 008,235 filed on 5 June 2014, and U.S. Serial No. 62 / 032,785 filed on 4 August 2014, the entire contents of which are incorporated herein by reference.
[0408] In some embodiments, the measurement method includes measuring the number of reads based on DNA sequences that map one or more predetermined chromosomes or chromosomal segments. Some such methods include creating a baseline value (cutoff value) for the number of DNA sequence reads that map to a particular chromosome or chromosomal segment, wherein a number of reads exceeding the baseline value indicates a particular genetic abnormality.
[0409] In some embodiments, the total measured amount of all alleles from one or more loci (e.g., the total amount from polymorphic or non-polymorphic loci) is compared to a reference amount. In some embodiments, the reference amount is (i) a threshold or (ii) an expected amount for a particular copy number hypothesis. In some embodiments, the reference amount (if no CNVs are present) is the total measured amount of all alleles from one or more loci for one or more chromosomes or chromosome segments that are known or expected to be free of deletions or duplications. In some embodiments, the reference amount (if CNVs are present) is the total measured amount of all alleles from one or more loci for one or more chromosomes or chromosome segments that are known or expected to have deletions or duplications. In some embodiments, the reference amount is the total measured amount of all alleles from one or more loci for one or more reference chromosomes or chromosome segments. In some embodiments, the reference amount is the mean or median of values determined for two or more different chromosomes, chromosome segments, or different samples. In some embodiments, the quantity of one or more polymorphic or non-polymorphic loci is determined using random (e.g., large-scale parallel shotgun sequencing) or labeled sequencing.
[0410] In some embodiments, the method includes using a reference amount to (a) measure the amount of genetic material in the target chromosome or chromosomal segment, (b) compare the amount obtained in step (a) with the reference amount, and (c) identify the presence or absence of deletions or duplications based on the comparison.
[0411] In some embodiments, the method includes sequencing DNA or RNA derived from a sample using a reference chromosome or chromosome segment to obtain a plurality of sequence tags to align with target loci. In some embodiments, the sequence tags are long enough to be assigned to specific target loci (e.g., 15 to 100 nucleotides in length). The target loci are derived from a plurality of different chromosomes or chromosome segments, including at least one first chromosome or chromosome segment suspected of having an abnormal distribution in the sample and at least one second chromosome or chromosome segment assumed to be normally distributed in the sample. In some embodiments, the plurality of sequence tags are assigned to the corresponding target loci. In some embodiments, the number of sequence tags to align with the target loci of the first chromosome or chromosome segment and the number of sequence tags to align with the target loci of the second chromosome or chromosome segment are determined. In some embodiments, these number of sequence tags are compared to determine whether the first chromosome or chromosome segment has an abnormal distribution (e.g., deletions or duplications).
[0412] In some embodiments, a value of f (e.g., fetal fraction or tumor fraction) is used to determine CNV, comparing the observed difference in the amount of two chromosomes or chromosome segments with the expected difference assuming a value of f for a particular type of CNV (e.g., U.S. Publications 2012 / 0190020, 2012 / 0190021, 2012 / 0190557, and 2012 / 0191358, the full contents of which are incorporated herein by reference). For example, as the fetal fraction increases, the difference between the amount of reference chromosome segments of the two chromosomes in the blood sample from the mother carrying the fetus and the amount of chromosome segments duplicated in the fetus also increases. Similarly, as the tumor fraction increases, the difference between the amount of reference chromosome segments of the two chromosomes and the amount of chromosome segments duplicated in the tumor also increases. In some embodiments, the method includes determining the likelihood of the CNV by comparing the relative frequency of the chromosome or chromosome segment of interest to a reference chromosome or chromosome segment (e.g., a chromosome or chromosome segment that is expected or known to be two chromosomes) with a value of f. For example, the difference in the amount of the first chromosome or chromosome segment and the reference chromosome or chromosome segment can be compared with what would be expected by assuming a value of f for each possible CNV (e.g., one or two extra copies of the chromosome segment of interest).
[0413] The following prediction examples illustrate how to distinguish between the duplication of the first homologous chromosome segment and the deletion of the second homologous chromosome segment using measurement / quantification methods. If the normal genomes of the two host chromosomes are considered as a baseline, the average difference between the baseline and cancer DNA in a mixture of normal and cancer cells can be obtained by analyzing the mixture. For example, suppose 10% of the DNA in a sample comes from cells with a deletion in the target chromosome region of the assay. In some embodiments, quantitative methods can be used to determine that the amount of reads corresponding to that region is expected to be 95% of the amount expected in a normal sample. This is because in each tumor cell with a deletion in the target region, one of the two target chromosome regions is missing, and the total amount of DNA mapped to that region is {90% (for normal cells) + (1 / 2 × 10%) (for tumor cells) = 95%}. Alternatively, in some embodiments, allele methods can be used to determine that the ratio of alleles present at heterozygous loci is, on average, 19:20. Here, suppose 10% of the DNA in a sample comes from cells that have undergone 5-fold local amplification of the target chromosome region of the assay. In some embodiments, quantitative methods reveal that the amount of reads corresponding to that region is expected to be 125% of the amount expected in a normal sample. This is because, in each tumor cell with 5x local amplification, one of the two target chromosomal regions is copied five extra times across that region, and the total amount of DNA mapped to that region is {90% (in normal cells) + ((2+5)×10% / 2) (in tumor cells) = 125%}. In addition, in some embodiments, allele methods reveal that the ratio of alleles present at heterozygous loci is 25:20 on average. Furthermore, using only the allele method, a chromosomal region in a sample with 10% cell-free DNA that has been locally amplified 5x appears the same as a deletion in the same region in a sample with 40% cell-free DNA. In these two cases, the haplotype that is under-occurring in the case of deletion is thought to be the haplotype without CNV in the case of local duplication. Conversely, the haplotype without CNV in the case of deletion is thought to be the haplotype that is over-occurring in the case of local duplication. By combining the likelihoods generated by this allele method with those generated by the quantitative method, two possibilities can be distinguished.
[0414] Examples of measurement / quantification methods using reference samples Examples of quantitative methods using one or more reference samples are described in U.S. Serial No. 62 / 008,235, filed on 5 June 2014, and U.S. Serial No. 62 / 032,785, filed on 4 August 2014, which are incorporated herein by reference in their entirety. In some embodiments, one or more reference samples are identified that are most likely to be free of CNVs in one or more chromosomes or the chromosome of interest (e.g., a normal sample) by selecting a sample with the highest fraction of tumor DNA, selecting a sample with the closest Z-value to zero, selecting a sample in which the data fits the hypothesis corresponding to the absence of CNVs with the highest confidence or likelihood, selecting a sample known to be normal, selecting a sample from an individual with the lowest likelihood of having cancer (e.g., young age, male in breast cancer screening, no family history, etc.), selecting a sample with the highest amount of DNA input, selecting a sample with the highest signal relative to the noise ratio, selecting a sample based on other criteria that appear to correlate with the likelihood of having cancer, or selecting a sample using a combination of several criteria. Once a reference set is selected, assuming these cases are dichromosome-dependent, the bias per SNP, i.e., bias due to experiment-specific amplification and other processing, can be estimated for each locus. This experiment-specific bias estimate can then be used to correct for bias in measurements of the target chromosome (e.g., the locus of chromosome 21), and, if necessary, other chromosomal loci, or samples that are not part of the subset assumed to be dichromosome 21. Once the bias is corrected for these samples with unknown ploidy, a second analysis of the data from these samples can be performed using the same or different methods to determine whether the individual (e.g., a fetus) has trichromosomal 21. For example, quantitative methods can be used on the remaining samples with unknown ploidy, and the Z-value can be calculated using the corrected measured gene data for chromosome 21. ...
Claims
1. A method for detecting a target tumor-specific mutation and monitoring the progression or state of cancer, (a) The tumor biopsy sample is sequenced to identify at least 10 tumor-specific mutations, where each tumor-specific mutation includes one or more single nucleotide variant (SNV) mutations. (b) Evaluate the sequencing results of the tumor biopsy sample and select at least 10 target loci, where each target locus is located such that it overlaps with at least one of the tumor-specific mutations identified for the tumor biopsy sample. (c) The cancer is monitored by assaying cell-free DNA isolated from multiple biological samples obtained from the subject at multiple different time points, wherein the assay includes performing targeted multiplex PCR amplification to amplify at least 10 target gene loci within a single reaction volume, wherein the single reaction volume is induced from the isolated cell-free DNA or DNA derived therefrom using PCR primers specific to the target gene loci, and then performing high-process sequencing on the amplified target gene loci to obtain sequence reads. Here, SNV mutations present in 0.015% or less of the cell-free DNA containing the SNV gene locus are detected from the sequence reads. Here, the biological sample is blood, serum, plasma, or urine sample. A method in which the presence of cell-free DNA containing one or more of the tumor-specific mutations in one or more of the biological samples indicates the progression or state of cancer.
2. The method according to claim 1, wherein the cell-free DNA includes circulating tumor DNA.
3. The method according to claim 1, wherein the tumor-specific mutation comprises one or more clonal SNV mutations.
4. The method according to claim 1, wherein the tumor-specific mutation comprises one or more subclonal SNV mutations.
5. The method according to claim 1, wherein the tumor-specific mutation comprises one or more clonal SNV mutations and one or more subclonal SNV mutations.
6. The method according to claim 1, wherein the tumor-specific mutation further comprises one or more copy number variation (CNV) mutations.
7. The method according to claim 1, further comprising determining the clonal heterogeneity of the tumor biopsy sample.
8. The method according to claim 1, further comprising designing a targeted PCR assay for the tumor-specific mutation identified in the tumor biopsy sample.
9. Step (b) includes evaluating the results of the sequencing of the tumor biopsy sample and selecting at least 50 target gene loci specific to the subject, wherein each target gene locus overlaps with at least one of the tumor-specific mutations, The method according to claim 1, wherein step (c) includes performing targeted multiplex PCR amplification to amplify the at least 50 target gene loci together in a single reaction volume, wherein the single reaction volume is induced from the isolated cell-free DNA or DNA derived therefrom using PCR primers specific to the target gene loci.
10. The method according to claim 1, wherein the cancer is colorectal cancer, lung cancer, bladder cancer, or breast cancer.
11. The method according to claim 1, further comprising detecting the recurrence and / or metastasis of the cancer from the tumor-specific mutation detected in the cell-free DNA.
12. The method according to claim 1, wherein step (a) includes performing whole exon sequencing on the target tumor biopsy sample and identifying the tumor-specific mutation.
Citation Information
Patent Citations
Genetic mutations associated with tumors
JP2010509922A
Analytical methods for cell free nucleic acids and applications
WO2012028746A1
Mutational analysis of plasma DNA for cancer detection
WO2013190441A2
Method for identifying novel minor histocompatibility antigens
WO2014026277A1
Systems and methods to detect rare mutations and copy number variation
WO2014039556A1