Detection of genetic or molecular aberrations associated with cancer

By analyzing chromosomal imbalances in cell-free DNA fragments, the accuracy and safety issues of existing cancer screening and monitoring technologies have been resolved, enabling efficient cancer diagnosis and prognostic monitoring.

CN114678128BActive Publication Date: 2026-01-09THE CHINESE UNIVERSITY OF HONG KONG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210291074.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2011-08-31
Filing Date
2011-11-30
Publication Date
2026-01-09
Estimated Expiration
2031-11-30

AI Technical Summary

Technical Problem

Existing cancer screening technologies have low accuracy and sensitivity, and imaging technologies may expose patients to radiation, making them ineffective for cancer screening, prognosis, and monitoring.

Method used

By analyzing cell-free DNA fragments, we can identify chromosomal imbalances caused by chromosomal fragment deletions and amplifications, and use computer systems to calculate changes in the properties of chromosomal regions to detect cancer-related genetic abnormalities.

Benefits of technology

It provides more efficient and accurate cancer diagnosis, screening, and prognostic monitoring, reduces the radiation risk to patients, and can track imbalances in chromosomal regions at different time points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114678128B_ABST
    Figure CN114678128B_ABST
Patent Text Reader

Abstract

The present invention provides systems, instruments and methods for determining genetic or molecular aberrations in biological samples from an organism. Biological samples, including cell-free DNA fragments, are analyzed to identify imbalances present in chromosomal regions, for example, due to tumor chromosomal deletions and / or amplifications. Multiple loci are used for the analysis of each chromosomal region. Such imbalances can then be used to diagnose (screen for) cancer and to prognose cancer patients, or to detect or monitor changes in a patient's health status prior to deterioration. Diagnoses, screens, prognoses and monitoring can be provided using the severity of genomic imbalances and the number of imbalanced regions. Systematic analysis of non-overlapping chromosomal segments can provide a general cancer screening approach. Furthermore, a patient can be tested at different time points to track the severity of each of one or more chromosomal regions and the number of chromosomal regions exhibiting chromosomal aberrations, thereby enabling screening, prognosing and monitoring of cancer progression (e.g., following treatment).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 61 / 418,391, filed November 30, 2010, entitled “DETECTION OF GENETIC ABERRATIONS ASSOCIATED WITH CANCER,” and U.S. Provisional Application No. 61 / 529,877, filed August 31, 2011, entitled “DETECTION OF GENETIC OR MOLECULAR ABERRATIONS ASSOCIATED WITH CANCER,” and is a non-provisional application thereof, the entire contents of which are incorporated herein by reference for all purposes.

[0003] This application relates to two jointly owned U.S. patent applications filed November 5, 2010, entitled “Size-Based Genomic Analysis” (U.S. Publication 2011 / 0276277) (Lo et al.) (Attorney’s File No. 80015-794101 / 006610US) and U.S. patent application filed November 5, 2010, entitled “Fetal Genomic Analysis From A Maternal Biological Sample” (U.S. Publication 2011 / 0105353) (Lo et al.) (Attorney’s File No. 80015-794103 / 006710US), the disclosures of which are incorporated herein by reference in their entirety. Background Technology

[0004] Cancer is a common disease affecting many people. It is often not discovered until severe symptoms appear. While some cancer screening technologies exist to determine the likelihood of a patient having cancer for common types, their reliability and accuracy are often low, or they require patients to be exposed to high doses of radiation. For many other types of cancer, there are currently no effective screening technologies.

[0005] Loss of heterozygosity (LOH) at specific loci has been detected in the peripheral blood DNA of patients with lung cancer and head and neck cancer (Chen XQ, et al. Nat Med 1996; 2:1033-5; Nawroz H, et al. Nat Med 1996; 2:1035-7). However, the relative small amount of LOH that can be detected has hindered the use of such techniques for detecting specific loci. Even with digital PCR, these methods cannot detect a relatively small amount of LOH. Furthermore, such techniques are still limited to studying known specific loci that occur in specific cancer types. Thus, such methods cannot or cannot effectively be used as a general means to screen for various cancers.

[0006] In addition to the limitations in screening for the presence of cancer, current techniques are not effective in providing prognostic diagnosis and monitoring of treatment effectiveness (e.g., recovery after surgery, chemotherapy, or immunotherapy or targeted therapy) for cancer patients. Because such techniques are often expensive (e.g., imaging techniques), inaccurate, inefficient, and insensitive, and if imaging techniques are used, can expose patients to radiation.

[0007] Thus, it would be desirable to have new techniques that can effectively provide screening, prognosis, and monitoring for cancer patients. SUMMARY

[0008] The present invention provides systems, apparatuses, and methods for detecting genetic aberrations associated with cancer in various embodiments. Biological samples, including cell-free DNA fragments, are analyzed to identify imbalances in chromosomal regions, e.g., due to loss and / or gain of chromosomal fragments, in tumors. Higher efficiency and / or accuracy can be achieved using chromosomal regions having multiple loci. Such imbalances can be used for diagnosis or screening of potential cancer patients and prognosis of cancer patients. Diagnosis, screening, prognosis, and monitoring of cancer can be provided using the severity of the imbalances and the number of imbalances. Furthermore, patients can be tested at different time points to track the degree of imbalances and the number of imbalances in each of one or more or a plurality of chromosomal regions, thereby enabling screening and prognosis, and monitoring (e.g., after treatment) of cancer.

[0009] According to one embodiment, a method is provided for analyzing a biological sample of an organism for deletions or amplifications of chromosomes associated with cancer. The biological sample includes nucleic acid molecules derived from normal cells as well as potentially cancer-associated cells. At least some of the nucleic acid molecules in the sample are free. First and second haplotypes are determined for the normal cells of the organism at a first chromosomal region. The first chromosomal region includes a first plurality of heterozygous loci. Each of the plurality of nucleic acid molecules in the sample has identifiable position information and respective allelic information in a reference genome of the organism in question. The position and the determined allelic information are used to determine a first set of nucleic acid molecules derived from the first haplotype and a second set of nucleic acid molecules derived from the second haplotype. A computer system is used to calculate a first value corresponding to the first set of nucleic acids and a second value corresponding to the second set of nucleic acids. Each value defines a property of the respective set of nucleic acid molecules (e.g., average size or number of molecules in the respective set). The first value is compared to the second value to determine a classification of whether the first chromosomal region exhibits a deletion or amplification in any of the cancer-associated cells.

[0010] According to another embodiment, a method is provided for analyzing a biological sample of an organism. The biological sample includes nucleic acid molecules derived from normal cells as well as potentially cancer-associated cells. At least some of the nucleic acid molecules in the sample are free. A plurality of non-overlapping chromosomal regions of the organism are identified. Each chromosomal region includes a plurality of loci. Each of the plurality of nucleic acid molecules in the sample has corresponding position information in a reference genome of the organism in question. For each chromosomal region, a respective set of nucleic acid molecules corresponding to the chromosomal region is identified based on the identified positions. The respective set of nucleic acid molecules includes at least one nucleic acid molecule falling on each of the plurality of loci of the chromosomal region. A computer system is used to calculate a respective value for each set, wherein each value defines a property of the nucleic acid molecules of the respective set. The respective values are compared to a reference value to determine whether the chromosomal region exhibits a deletion or amplification. Chromosomal regions classified as exhibiting a deletion or amplification of a set of chromosomes are then quantitatively determined.

[0011] According to another embodiment, there is provided a method for determining the progression of chromosomal aberrations in an organism using a biological sample, wherein the biological sample comprises nucleic acid molecules derived from normal cells as well as potentially cancer-related cells. In the biological sample, at least some of the nucleic acid molecules are free. One or more non-overlapping chromosomal regions are identified with respect to a reference genome of the organism. Each chromosomal region comprises a plurality of loci. Samples taken from the organism at different time points are analyzed to determine the progression of the pathology. With respect to a sample, each of the plurality of nucleic acid molecules in the sample has a location in the identified reference genome of the organism. With respect to each chromosomal region, based on the identified locations, a respective set of nucleic acid molecules is identified that corresponds to the chromosomal region. The respective set of nucleic acid molecules comprises at least one nucleic acid molecule at each of the plurality of loci of the chromosomal region. A computer system computes a respective value for each set of nucleic acid molecules. The respective value defines a property of the respective set of nucleic acid molecules. The respective value is compared to a reference value to determine whether a deletion or amplification is present in the first chromosomal region. The progression of chromosomal aberrations in the organism is then determined using the possible deletions or amplifications in the plurality of chromosomal regions at the plurality of time points.

[0012] Other embodiments of the present application are directed to systems, portable devices, and computer readable media associated with the methods described herein.

[0013] The features and advantages of the present application will be better understood with reference to the following detailed description when considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 Chromosomal regions in which deletional aberrations are present in cancer cells are described.

[0015] Figure 2 Chromosomal regions in which amplificational aberrations are present in cancer cells are described.

[0016] Figure 3 Table 300 shows different types of cancer, the associated regions, and their corresponding aberrations.

[0017] Figure 4 Described are embodiments according to the present application in which the dosage corresponding to chromosomal regions that do not exhibit aberrations in cancer cells is measured in plasma.

[0018] Figure 5 Described are embodiments according to the present application in which the dosage corresponding to chromosomal regions 510 in which deletions are present in cancer cells is measured in plasma, thereby determining the corresponding deleted regions.

[0019] Figure 6 Described are embodiments according to the present application in which the dosage corresponding to chromosomal regions 610 in which amplifications are present in cancer cells is measured in plasma, thereby determining the corresponding amplified regions.

[0020] Figure 7 An RHDO analysis of plasma DNA from a patient with hepatocellular carcinoma (HCC) is described according to embodiments of the application for a chromosomal segment located on chromosome lp, wherein the chromosomal segment shows a single allele amplification in tumor tissue.

[0021] Figure 8 Changes in the length distribution of nucleic acid fragments in plasma corresponding to each of the two haplotypes of the chromosomal region are described according to embodiments of the application when the tumor has a chromosomal deletion.

[0022] Figure 9 Changes in the length distribution of nucleic acid fragments in plasma corresponding to each of the two haplotypes of the chromosomal region are described according to embodiments of the application when the tumor has a chromosomal deletion.

[0023] Figure 10 A flowchart of a method for analyzing haplotypes of a biological sample of an organism to determine whether a chromosomal region exhibits a deletion or an amplification according to embodiments of the application is described.

[0024] Figure 11 A region 1110 with a deletion and a sub-region 1130 in a cancer cell and a quantitative measurement of the region in plasma to determine the region with a deletion according to embodiments of the application are described.

[0025] Figure 12 How the location of the aberration can be mapped using RHDO analysis according to embodiments of the application is described.

[0026] Figure 13 A classification of RHDO starting from another direction according to embodiments of the application is described.

[0027] Figure 14 A flowchart of a method 1400 for analyzing a biological sample of an organism using multiple chromosomal regions according to embodiments of the application is provided.

[0028] Figure 15 A table 1500 illustrating the sequencing depth required for different numbers of aberrant fragments at different tumor nucleic acid relative percentage concentrations according to embodiments of the application is described. Figure 15 An evaluation of the number of molecules to be tested performed at different tumor nucleic acid relative percentage concentrations in a sample is provided.

[0029] Figure 16The principles of measuring the relative percentage concentration of tumor nucleic acids in plasma by relative haplotype dosage (RHDO) analysis are illustrated according to embodiments of the application. Hap I and Hap II represent two haplotypes in non-tumor tissue according to embodiments of the application.

[0030] Figure 17 A flow chart illustrating a method of using a biological sample comprising nucleic acid molecules to determine the progression of chromosomal aberrations in an organism according to embodiments of the application is illustrated.

[0031] Figure 18A SPRT curves are shown for RHDO analysis of a chromosomal segment on the q arm of chromosome 4 in a patient with cancer. The circle points represent the ratio of the cumulative counts of all analyzed sites upstream of each heterozygous locus. Figure 18B SPRT curves are shown for RHDO analysis of a chromosomal segment on the q arm of chromosome 4 in a patient after treatment.

[0032] Figure 19 Common chromosomal aberrations found in HCC are shown.

[0033] Figure 20A Results obtained from analysis of HCC patients and healthy control subjects using targeted analysis (e.g., targeted enrichment capture sequencing) are shown, normalized to the proportion of sequencing reads. Figure 20B Results obtained from analysis of nucleic acid length in plasma after targeted enrichment capture sequencing of 3 HCC patients and 4 healthy control subjects are shown.

[0034] Figure 21 A Circos plot of an HCC patient obtained according to embodiments of the application is shown, depicting data obtained from sequencing tag counts of plasma DNA.

[0035] Figure 22 Sequencing tag count analysis of plasma samples from chronic hepatitis B virus (HBV) carriers who do not have HCC is shown according to embodiments of the application.

[0036] Figure 23 Sequencing tag count analysis of plasma samples from a patient with stage III nasopharyngeal carcinoma (NPC) is shown according to embodiments of the application.

[0037] Figure 24 Sequencing tag count analysis of plasma samples from a patient with stage IV NPC is shown according to embodiments of the application.

[0038] Figure 25A plot of the cumulative frequency of plasma DNA length against regions in which loss of heterozygosity (LOH) is present in tumor tissue is shown, according to embodiments of the application.

[0039] Figure 26 A plot of the relationship between AQ and plasma DNA length in regions of LOH is shown. According to embodiments of the application, AQ reaches 0.2 at a length of 130 bp.

[0040] Figure 27 A plot of the cumulative frequency of plasma DNA length in regions of chromosome amplification present in tumor tissue is shown, according to embodiments of the application.

[0041] Figure 28 A plot of the relationship between AQ and plasma DNA length for amplified regions is shown, according to embodiments of the application.

[0042] Figure 29 A block diagram of an example of a computer system 900 usable by systems and methods according to embodiments of the application is shown.

[0043] Definitions

[0044] As used herein, the term "biological sample" refers to any sample taken from a subject (e.g., a human, a human with cancer, a human suspected of having cancer, or other organism) and comprising one or more nucleic acid molecules of interest.

[0045] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in either single- or double-stranded form, and polymers thereof. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in the same manner. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), copy number variants, complementary sequences, and sequences that are substantially identical thereto. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid encompasses, but is not limited to, genes, complementary DNA (cDNA), messenger RNA (mRNA), small molecule non-coding RNA, microRNA (miRNA), Piwi-interacting RNA, and short hairpin RNA (shRNA) or other sequences encoded by a gene or locus or other sequences on a chromosome.

[0046] The term "gene" refers to a segment of DNA involved in producing a polypeptide chain or transcribed RNA product. It can include regions preceding and following the coding region (leader and trailer), as well as intervening sequences (introns) between individual coding segments (exons).

[0047] As used herein, the term "clinically relevant nucleic acid sequence" or "clinically relevant chromosomal region" (or region / segment to be interrogated) can refer to a polynucleotide sequence corresponding to a larger genomic sequence segment for which there is a potential imbalance to be interrogated or to the larger genomic sequence itself. Examples include a genomic segment that is deleted or amplified, or potentially deleted or amplified (including simple duplication), or a larger region that includes a sub-region of the segment. In some embodiments, multiple clinically relevant nucleic acid sequences, or multiple equivalent markers of a clinically relevant nucleic acid sequence, can provide data for detecting an imbalance in the region. For example, data from five non-contiguous sequences on a chromosome can be used in an additive manner to determine a possible imbalance, effectively reducing the dosage of the sample of interest to 1 / 5.

[0048] As used herein, the term "reference nucleic acid sequence" or "reference chromosomal region" refers to a nucleic acid sequence whose dosage profile or length profile is compared to that of a test region. Examples of reference nucleic acid sequences include a chromosomal region that does not contain a deletion or amplification, a complete genome (e.g., normalized by sequencing tag counts), a region derived from one or more samples known to be normal (which can be the same region as the sample to be tested), or a chromosomal region of a particular haplotype. Such reference nucleic acid sequences can be present endogenously in the sample or added exogenously during sample processing or analysis. In some embodiments, the reference chromosomal region demonstrates a length profile representative of a healthy state not suffering from a disease. In other embodiments, the reference chromosomal region demonstrates a quantitative profile representative of a healthy state not suffering from a disease.

[0049] As used herein, the term "based on" means "based, at least in part, on" and refers to one or more values (or results) used in the determination of another value, e.g., formed in the relationship of an input to a method and an output of the method. As used herein, the term "derive" also refers to the relationship of an input to a method and an output of the method, e.g., formed in the calculation of a formula.

[0050] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative data set and / or a quantitative relationship between data sets. For example, the ratio (or function of the ratio) of a first quantity of a first nucleic acid sequence to a second quantity of a second nucleic acid sequence is a parameter.

[0051] As used herein, the term "locus" is the position or address of any length of nucleotides (or base pairs) that can vary in a genome.

[0052] As used herein, the term "sequence imbalance" or "aberration" refers to any significant deviation from a defined at least one threshold of a reference amount of a clinically relevant chromosomal region. Sequence imbalances can include chromosomal dosage imbalances, allelic imbalances, mutation dosage imbalances, copy number imbalances, haplotype dosage imbalances, and other similar imbalances. For example, an allelic imbalance can be formed due to a tumor having one deleted allele of a gene or one amplified allele of a gene, or differential amplification of two alleles in its genome, thereby forming an imbalance at a particular locus of the sample. As another example, a patient can have a genetic mutation in a tumor suppressor gene. The patient can then go on to develop a tumor in which the non-mutated allele of the tumor suppressor gene is deleted. Thus, in the tumor, there is a mutation dosage imbalance. When the tumor releases its DNA into the patient's plasma, the tumor DNA will mix with the patient's normal somatic structural DNA in the plasma. Using the methods described herein, the mutation dosage imbalance present in this DNA mixture in the plasma can be detected.

[0053] As used herein, the term "haplotype" refers to a combination of alleles at multiple loci, where the multiple alleles are transmitted together to offspring on the same chromosome or chromosomal region. A haplotype can refer to as few as a pair of loci or chromosomal regions, or to an entire chromosome. The term "allele" refers to an alternative DNA sequence at the same physical location on the same chromosome, which can result in the same or different phenotypic characteristics. In any particular diploid organism, whose chromosomes have two copies (except for the sex chromosomes in male subjects), the genotype of each gene includes the paired alleles present at that locus, which are identical if homozygous, and different if heterozygous. A population of organisms or species typically includes multiple alleles at each locus in different individuals. A locus in which more than one allele is found in a population is referred to as a polymorphic site. The extent of variation in alleles at a locus can be measured in terms of the number of alleles (i.e., the extent of polymorphism) or the proportion of heterozygotes in the population (i.e., the proportion of heterozygosity). As used herein, the term "polymorphism" refers to any inter-individual variation in a human gene, regardless of its frequency. Examples of such variations include, but are not limited to, single nucleotide polymorphisms, simple tandem repeat polymorphisms, insertion-deletion polymorphisms, mutations (which can be the cause of disease), and changes in copy number.

[0054] The term "sequencing tag" refers to a sequence determined from all or part of a nucleic acid molecule (e.g., a DNA fragment). Typically, only one end of the fragment is sequenced, e.g., about 30 bp. The sequencing tag is then aligned to a reference genome. Alternatively, both ends of the fragment can be sequenced, generating two sequencing tags, which can provide higher alignment accuracy, and can also provide fragment length information.

[0055] The term "universal sequencing" refers to the fact that primers for sequencing, which are attached to the ends of fragments for sequencing, can be complementary to the adapter sequence for sequencing. Thus, any fragment can be sequenced using the same primers, and thus sequencing can be random.

[0056] The term "size distribution" refers to any one value or set of values representing the length, mass, weight, or other measure of size of molecules corresponding to a particular group (e.g., fragments resulting from a particular haplotype or a particular chromosomal region). Various embodiments can use various size distributions. In some embodiments, the size distribution relates to the relative order of lengths of one chromosomal-related fragment relative to other chromosomal-related fragments (e.g., mean, median, or geometric mean). In other embodiments, the size distribution can relate to a statistical value of the actual size of the chromosomal fragments. In one implementation, the statistical value can include any mean, geometric mean, or median of the chromosomal fragments. In another implementation, the statistical value can include the total length of fragments below a certain threshold, which can be divided by the total length of all fragments or at least the total length of fragments below a certain larger cutoff value.

[0057] As used herein, the term "classification" refers to any quantity or other characteristic related to a particular property of a sample. For example, a "+" symbol (or the word "positive") can indicate that a sample is classified as having a chromosomal aberration of a deletion or amplification. The classification can be binary (e.g., positive or negative), or have a higher level of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" refer to a predetermined value used in an operation. For example, a cutoff size can refer to fragments above which value will be excluded. A threshold can be a value above or below which a particular classification can be applied. Either of these terms can be used in any of these contexts.

[0058] The term "level of cancer" can refer to whether cancer is present, the stage of cancer, the size of a tumor, the degree of deletion or amplification involving a chromosomal region (e.g., double duplication amplification or triple duplication amplification), and / or other measures of severity of cancer. The level of cancer can be a quantity or other characteristic. The level of cancer can be zero. In addition, the level of cancer also includes pre-malignant or pre-cancerous states related to deletions or amplifications. DETAILED DESCRIPTION

[0059] Cancerous tissue (tumor) can have aberrations, such as deletions or amplifications of chromosomal regions. The tumor can release DNA fragments into the fluidic tissue of the body, such as peripheral blood. Various embodiments can identify the tumor by analyzing the DNA fragments relative to the normal (expected) value of DNA in the chromosomal region, thereby identifying the aberration.

[0060] The exact size and location of the deletions or amplifications can vary. Perhaps there are known aberrations common to cancer or to a particular cancer in a particular chromosomal region (from which a particular cancer can be diagnosed). When the particular region is unknown, a systematic approach for analyzing the entire genome or a large portion of the genome can be used to detect aberrant regions, which can be scattered throughout the genome and vary in size (e.g., the number of bases deleted or amplified). Chromosomal regions can be tracked at different time points to detect changes in the severity of an aberration or changes in the number of aberrant regions. Such tracking can provide important information for cancer screening, prognostic diagnosis, and monitoring of a tumor (e.g., after treatment or for detecting recurrence or tumor progression).

[0061] The detailed description of the invention begins with examples of chromosomal aberrations in cancer. Next, examples of ways to detect chromosomal aberrations by detecting and analyzing cell-free DNA in a biological sample are discussed. Once a method for detecting aberrations in a chromosomal region is established, methods for using detection of aberrations in many chromosomal regions in a systematic manner for cancer screening (diagnosis) and prognosis of a patient are further described. In addition, the detailed description of the invention describes methods for tracking cancer-related quantitative indicators obtained from testing chromosomal aberrations by detecting one or more regions at multiple time points to provide screening, prognosis, and monitoring of a patient. Examples are then discussed.

[0062] I. Examples of chromosomal aberrations in cancer

[0063] Chromosomal aberrations are ubiquitous in cancer cells. In addition, specific cancers have characteristic patterns of specific chromosomal aberrations. For example, gains in the amount of DNA corresponding to arms 1p, 1q, 7q, 15q, 16p, 17q, and 20q and losses in the amount of DNA corresponding to 3p, 4q, 9p, and 11q are commonly detected in hepatocellular carcinoma (HCC). Previous studies have demonstrated that such genetic aberrations can also be detected in the peripheral blood DNA of cancer patients. For example, loss of heterozygosity (LOH) at specific loci has been reported in the peripheral blood DNA molecules of patients with lung cancer and head and neck cancer (Chen XQ, et al. Nat Med 1996; 2:1033-5; Nawroz H, et al. Nat Med 1996; 2:1035-7). The genetic aberrations detected in the plasma or serum are consistent with those found in tumor tissue. However, because tumor-derived DNA constitutes only a small fraction of the total circulating DNA in the peripheral blood, allelic imbalance resulting from LOH in tumor cells is usually weak. Many investigators have developed digital polymerase chain reaction (PCR) techniques (Vogelstein B, Kinzler KW. Proc Natl Acad Sci U S A. 1999; 96:9236-41; Zhou W, et al. Nat Biotechnol 2001; 19:78-81; Zhou W, et al. Lancet. 2002; 359:219-25) for accurate quantification of different alleles at a locus in peripheral blood DNA molecules (Chang HW, et al. J Natl Cancer Inst. 2002; 94:1697-703). Digital PCR is more sensitive than real-time PCR or other DNA quantification methods for detecting weak allelic imbalance resulting from LOH at specific loci in tumor DNA. However, digital PCR still has difficulty in discriminating very weak allelic imbalance at specific loci, and thus the embodiments described herein analyze different regions of the chromosome in an integrated manner.

[0064] In addition, the techniques described herein can be applied to detect a pre-malignant or pre-cancerous state. Examples of such states include cirrhosis of the liver and cervical intraepithelial neoplasia. The state of cirrhosis of the liver is a pre-malignant state for liver cancer, and the state of cervical intraepithelial neoplasia is a pre-malignant state for cervical cancer. Such pre-malignant states are reported to have a number of molecular alterations as they evolve into malignant tumors. For example, the presence of LOH at chromosome arms 1p, 4q, 13q, 18q, and simultaneous deletion at more than 3 loci is associated with an increased risk of HCC in patients with cirrhosis (Roncalli M et al. Hepatology 2000; 31 :846-50). In addition, such pre-malignant lesions also release DNA into the peripheral blood circulation, but possibly at lower concentrations. The techniques described can detect deletions or amplifications by analyzing the DNA fragments in the plasma and measuring the concentration (including the relative percentage concentration) of pre-malignant cell-free DNA in the plasma. Such aberrations can be readily detected (e.g., the depth of sequencing or the number of such changes detected), and the concentration can predict the likelihood or speed of progression to a fully blown cancerous state.

[0065] A. Deletion of a Chromosomal Region

[0066] Figure 1 A chromosomal region in which a deletion aberration exists in a cancer cell is set forth. A normal cell shows two haplotypes, Hap I and Hap II. As shown in Figure 1 each of a plurality of heterozygous loci 110 (also referred to as single nucleotide polymorphisms, SNPs), Hap I and Hap II have sequence information. In a cell associated with cancer, Hap II has a deleted chromosomal region 120. For example, a cell associated with cancer can be obtained from a tumor (e.g., a malignant tumor), from a metastatic lesion of a tumor (e.g., in a regional lymph node or in a distant organ), or from a pre-cancerous or pre-malignant lesion, such as those described above.

[0067] In the chromosomal region 120 of a cancer cell, in which one of the two homologous haplotypes is deleted, all of the heterozygous SNPs 110 appear homozygous due to the absence of the other corresponding allele on the deleted homologous chromosome. Thus, this type of chromosomal aberration is referred to as loss of heterozygosity (LOH). In the region 120, the non-deleted allele of these SNPs represents one of the two haplotypes that can be found in normal tissue. In the region 120, the deleted allele of these SNPs represents the other haplotype that is not found in normal tissue. Figure 1In the example shown, the haplotype at the LOH region 120, Hap I, can be determined by the genotype of the tumor tissue. The other haplotype, Hap II, can be determined by comparing the genotype of the normal tissue to the genotype on the surface of the cancer tissue. Hap II can be constructed by linking all of the missing alleles. That is, the alleles that are only present in the normal cells at region 120 (but which are missing in the cancer cells at region 120) are determined to be the same haplotype, Hap I. Through this analysis, the haplotype of the patient (e.g., a hepatocellular carcinoma, HCC, patient) that corresponds to all of the chromosomal regions that exhibit LOH in the tumor tissue can be determined. Such a method is useful only for cancer cell analysis and is effective only for determining the haplotype in region 120, but it provides a good illustration of the missing regions in the chromosome.

[0068] B. Amplification of Chromosomal Regions

[0069] Figure 2 Cancer cells are shown to exhibit chromosomal regions that exhibit amplification aberrations. Normal cells show two haplotypes, Hap I and Hap II. As shown in the example, both Hap I and Hap II have sequence information in each of the multiple heterozygous loci 210. In the tumor cells, Hap II has a region of the chromosome 220 that is amplified by a factor of two (duplicated). Figure 2

[0070] Similarly, for regions in the tumor tissue that have single allelic amplification, the amplified alleles at the SNPs 210 can be detected by methods such as chip analysis. One of the two haplotypes can be determined by linking all of the amplified alleles in the chromosomal region 220 (e.g., Hap II in the example shown). The amplified alleles at a particular locus can be determined by comparing the number of alleles at the locus. Then, the other haplotype (Hap I) can be determined by linking the non-amplified alleles. Such a method is useful only for cancer cell analysis and is effective only for determining the haplotype in region 220, but it provides a good illustration of the amplified regions in the chromosome. Figure 2

[0071] Amplification can result from regions of more than two chromosomes, or from a duplication of a gene in one chromosome. A region can be duplicated in tandem, or a region can be a microchromosome that includes one or more copies of the region. In addition, amplification can also result from a gene of one chromosome being duplicated and the duplication product being inserted into a different chromosome or a different region of the same chromosome. Such insertions are a type of amplification.

[0072] II. Selection of Chromosomal Regions ​​

[0073] Genomic aberrations in cancerous tissue can be detected in samples such as plasma and serum when the cancerous tissue contributes at least a portion of cell-free DNA (and potentially intracellular DNA). The challenge in detecting these aberrations is that tumors or cancers can be quite small, resulting in relatively weak DNA contribution from cancer cells. Consequently, the amount of cell-free DNA with aberrations is relatively small, making detection extremely difficult. At a single locus in the genome containing the aberration to be detected, there may not be sufficient DNA. The method described herein overcomes this difficulty by analyzing DNA at chromosomal regions (haplotypes) comprising multiple loci, thereby clustering small changes at individual loci into perceptible differences based on haplotype. Therefore, analyzing multiple loci within a region provides greater accuracy and precision, and reduces false positives and false negatives.

[0074] Furthermore, the aberrant regions can be quite small, making them difficult to identify. If only one locus or specific loci are used, aberrations not located at those loci will be missed. As described in this paper, methods can be used to study entire regions, thereby discovering aberrations present in subregions within those regions. When the analyzed region covers the entire genome, the entire genome can be analyzed to discover aberrations of varying lengths and locations, as described in more detail below.

[0075] To illustrate these points, as shown above, regions can exhibit distortion. However, the region used for analysis must be carefully selected. The length and location of the region can alter the results and thus affect the analysis. For example, if the analysis... Figure 1 In the first region shown, no aberrations were detected. If the second region is analyzed, for example using the method described herein, aberrations can be detected. When analyzing a larger region that includes both the first and second regions, one encounters the problem that only a portion of the larger region may have aberrations, making it more difficult to identify any aberrations, and one faces the problem of determining the exact location and length of the aberration. Various embodiments of the invention can solve some and / or all of these problems. The description for selecting regions also applies to haplotypes using the same chromosomal region or haplotypes using two different chromosomal regions.

[0076] A. Selecting specific chromosomal regions

[0077] In one embodiment, the specific region can be selected based on the knowledge of the cancer or the patient. For example, the region is known to be ubiquitously aberrant in many cancers or in a specific cancer. The exact length and location of the corresponding region can be obtained based on reference to some well-known literature related to the type of cancer or specific risk factors the patient has. In addition, the patient's tumor tissue can be obtained and analyzed to identify the aberrant region, as described above. Currently, such techniques require the acquisition of cancer cells, which can not be practical for a patient just diagnosed, but such techniques can be used to identify regions to monitor in the same patient at different time points (e.g., after surgical removal of cancerous tissue, or after chemotherapy or immunotherapy or targeted therapy, or to detect tumor recurrence or progression).

[0078] One can identify more than one specific region. Each such region can be used for analysis independently, or different regions can be analyzed collectively. In addition, the regions can be subdivided to provide higher accuracy in locating aberrations.

[0079] Figure 3 A table 300 is shown to display different types of cancer, the relevant regions, and their corresponding aberrations. Column 310 lists different cancer types. The embodiments described herein can be used for any type of cancer related to aberrations, and thus the list is merely an example. Column 320 shows regions (e.g., large regions such as 7p or more specific regions such as 17q25) where gains (amplifications) are associated with the specific cancer of the same row. Column 330 shows regions where deletions (deletions) can be found. Column 340 lists references that discuss the relevance of these regions to specific cancers.

[0080] According to the methods described herein, these regions with potential chromosomal aberrations can be treated as candidate chromosomal regions for aberration analysis. Examples of other genomic regions that are altered in cancer can be found in the databases of the Cancer Genome Anatomy Project (cgap.nci.nih.gov / Chromosomes / RecurrentAberrations) and Atlas of Genetics and Cytogenetics in Oncology and Haematology (atlasgeneticsoncology.org Tumors / Tumorliste.html).

[0081] As one can see, the identified regions can be quite large, while other regions can be more specific. The aberration does not necessarily include the entire region identified in the table. Thus, for a particular patient, such clues of the type of aberration cannot be precisely pinpointed, but can be used more as a rough guide for analysis of large regions. Such large regions can include many sub-regions (which can be of equal size). These can be analyzed individually as well as collectively (details of which are described herein). Thus, depending on the particular circumstances of the cancer to be tested, multiple embodiments can combine aspects of selecting large regions, but can also use more general techniques as described below.

[0082] B. Selecting Arbitrary Chromosomal Regions

[0083] In another embodiment, the chromosomal regions to be analyzed are selected arbitrarily. For example, the genome can be divided into regions of 1 megabase (Mb) in length, or other predetermined piece lengths, such as 500 Kb or 2 Mb. If the regions are 1 Mb, since there are approximately 3 billion bases in the haploid human genome, there are approximately 3000 regions in the human genome. As discussed in more detail below, each of these regions can be analyzed.

[0084] The determination of such regions is not based on any knowledge of the cancer or patient, but rather on the systematic division of the genome into regions to be analyzed. In one implementation, when a chromosome does not have a length of a number of predetermined pieces (e.g., is not divisible by 1 million bases), then the last region of the chromosome can be less than the predetermined length (e.g., less than 1 Mb). In another implementation, each chromosome can be divided into regions of equal length (or approximately equal - within rounding error) depending on the total length of the chromosome and the number of pieces to be created (which is generally variable among chromosomes). In such an implementation, the length of the segments of each chromosome can vary.

[0085] As mentioned above, particular regions can be identified depending on the particular cancer to be tested, but the particular regions can also be subdivided into smaller regions (e.g., equal size sub-regions covering a larger particular region). In this manner, the aberration can be pinpointed. In the discussion below, any general reference to a chromosomal region can be to a specifically identified region and / or an arbitrarily selected region.

[0086] III. Detection of Aberrations in Particular Haplotypes

[0087] This section describes methods for detecting aberrations in a single chromosomal region by analyzing a biological sample containing cell-free DNA. In embodiments of this section, the single chromosomal region is a region containing multiple heterozygous (different alleles) loci, whereby two haplotypes can be distinguished by knowing the specific allele at a given locus. Thus, a given nucleic acid molecule (e.g., a fragment of cell-free DNA) can be identified as having come from a specific one of the two haplotypes. For example, a fragment can be sequenced, resulting in a sequence tag that aligns onto the chromosomal region, and then the haplotype at the heterozygous locus to which the allele belongs can be identified. Two general techniques for determining aberrations in a specific haplotype (Hap) are described below, specifically, tag counting and size analysis.

[0088] A. Determining haplotypes

[0089] To distinguish between the two haplotypes, the two haplotypes of the chromosomal region are first determined. For example, the two haplotypes Hap I and Hap II shown by the normal cell of Figure 1 Figure 1 In this embodiment, the haplotype includes a first plurality of loci 110 that are heterozygous and allow the two haplotypes to be distinguished. This first plurality of loci covers the chromosomal region to be analyzed. The alleles on the different heterozygous loci (hets) can first be determined, and then the haplotype of the patient can be determined.

[0090] The haplotype of the SNP alleles can be determined by single molecule analysis methods. Examples of such methods have been described by Fan et al. (Nat Biotechnol. 2011; 29: 51-7), Yang et al. (Proc Natl Acad Sci U S A. 2011; 108: 12-7), and Kitzman et al. (Nat Biotechnol. 2011 Jan; 29: 59-63). In addition, the haplotype of an individual can be determined by analyzing the genotypes of family members (e.g., parents, siblings, and children). Examples include methods described by Roach et al. (Am J Hum Genet. 2011; 89(3): 382-97) and Lo et al. (Sci Transl Med. 2010; 2: 61ra91). In another embodiment, the haplotype of an individual can be determined by comparing the genotyping results of the tumor tissue to the genotyping results of the normal structural genome. The genotypes of these subjects can be obtained by microarray analysis (e.g., using t).

[0091] ​In addition, haplotypes can also be constructed by other methods familiar to those skilled in the art. Examples of such methods include haplotype determination based on single molecule analysis, such as digital PCR (Ding C and Cantor CR. Proc Natl Acad Sci USA 2003; 100:7449-7453; Ruano G et al. Proc Natl Acad Sci USA 1990; 87:6296-6300), chromosome picking or separation (Yang H et al. Proc Natl Acad Sci U S A 2011; 108: 12-17; Fan HC et al. Nat Biotechnol 2011; 29:51-57), sperm haplotyping (Lien S et al. Curr Protoc Hum Genet 2002; Chapter 1: Unit 1.6), and imaging techniques (Xiao M et al. Hum Mutat 2007; 28:913-921). Other methods include haplotype analysis techniques based on allele-specific PCR (Michalatos-Beloin S et al. Nucleic Acids Res 1996; 24:4841-4843; Lo YMD et al. Nucleic Acids Res 19:3561-3567), cloning and restriction enzyme digestion (Smirnova AS et al. Immunogenetics 2007; 59:93-8), and the like. Other methods are based on the distribution and linkage disequilibrium structure of haplotype blocks in a population, which allows the haplotype of a subject to be derived from statistical evaluation (Clark AG. Mol Biol Evol 1990; 7: 111-22; 10: 13-9; Salem RM et al. Hum Genomics 2005; 2:39-66).

[0092] Another method of determining the haplotype of the region of LOH is by genotyping the normal tissue and tumor tissue (if tumor tissue is available) of the subject. In the presence of LOH, tumor tissue with a relatively high percentage concentration of tumor cells will show apparent homozygosity at all SNP loci within the region of LOH. The genotypes of these SNP loci include one haplotype (Hap I) shown below. Figure 1 On the other hand, in the normal tissue, the SNP loci within the region of LOH will show heterozygosity. The allele present in the normal tissue but not in the tumor tissue includes the other haplotype (Hap II) shown below. Figure 1Hap II) of the LOH region shown.

[0093] B. Relative haplotype dosage (RHDO) analysis

[0094] As mentioned above, a chromosomal aberration having amplification or deletion of one of the haplotypes of a chromosomal region results in a dosage imbalance of the two haplotypes in the corresponding chromosomal region in tumor tissue. In the plasma of a person having tumor growth, a fraction of the peripheral blood DNA is derived from tumor cells. Because of the presence of tumor-derived DNA in the plasma of cancer patients, such imbalances are also present in their plasma. The imbalance in dosage of the two haplotypes can be detected by counting the number of molecules derived from each haplotype.

[0095] In terms of a chromosomal region in which LOH is observed in tumor tissue (e.g. Figure 1 In terms of a chromosomal region in which copy number amplification is observed in tumor tissue, Hap II is relatively overrepresented when compared to Hap I in peripheral blood DNA molecules (fragments) due to the lack of contribution from Hap II of the tumor tissue. In terms of a chromosomal region in which copy number amplification is observed in tumor tissue, Hap II is relatively overrepresented when compared to Hap I in terms of regions affected by monoallelic amplification of Hap II due to the release of extra dosage of Hap II by the tumor tissue. To determine over- or under-representation, certain DNA fragments derived from Hap I or Hap II in a sample can be determined by various methods, such as methods based on universal sequencing and alignment analysis, or using digital PCR and sequence-specific probes, among others.

[0096] After sequencing a plurality of DNA fragments obtained from the plasma (or other biological sample) of a cancer patient to generate sequence tags, the sequence tags corresponding to alleles on both haplotypes can be identified and counted. The number of sequence tags corresponding to each of the two haplotypes is then compared to determine whether the two haplotypes are equally present in the plasma. In one embodiment, sequential probability ratio test (SPRT) can be used to determine whether there is a significant difference in the presence of the two haplotypes in the plasma. A statistically significant difference indicates the presence of a chromosomal aberration in the analyzed chromosomal region. Furthermore, the quantitative difference in the two haplotypes in the plasma can be used to estimate the relative percentage concentration of tumor-derived DNA in the plasma, as described below.

[0097] The diagnostic methods described herein for determining the identity of a DNA fragment (e.g., its location in the human genome) are not limited to the use of massively parallel sequencing as a detection platform according to embodiments of the application. In addition, these diagnostic methods can also be used, for example, but not limited to, microfluidic digital PCR systems (e.g., the Fluidigm Digital Array system), droplet digital PCR systems (e.g., RainDance and QuantaLife), BEAM-ing systems (i.e., a system based on Bead, Emulsion PCR, Amplification and Magnetism) (Diehl et al. Proc Natl Acad Sci USA 2005; 102: 16368-16373), real-time PCR, mass spectrometry-based systems (e.g., the Sequenom MassArray system), and multiplex ligation-dependent probe amplification (MLPA) analysis.

[0098] Normal region

[0099] Figure 4 A chromosome region within a cancer cell that does not exhibit aberrations and measurements made in plasma according to embodiments of the application are shown. The chromosome region 410 can be selected by any method, for example, a method based on a specific cancer to be tested or a method based on a general screening (i.e., one that uses predetermined fragments that cover a large portion of the genome). To distinguish between the two haplotypes, the two haplotypes are first determined. Figure 4 The two haplotypes (Hap I and Hap II) for a normal cell are shown with respect to the chromosome region 410. The haplotypes include a first plurality of loci 420. The first plurality of loci 420 spans the chromosome region 410 to be analyzed. As shown, the loci are heterozygous in the normal cell. The two haplotypes for a cancer cell are also shown. In the cancer cell, no regions are deleted or amplified.

[0100] In addition, Figure 4The number of allele counts on each haplotype for each locus 420 is also shown. In addition, the cumulative total for certain sub-regions of the chromosome region 410 is provided. The number of allele counts corresponds to the number of DNA fragments that correspond to a particular haplotype at a particular locus. For example, a DNA fragment that includes the first locus 421 and has allele A takes a count of Hap I. While a DNA fragment that has allele T takes a count of Hap II. The determination of where a fragment aligns (i.e., whether it includes a particular locus) and what allele it includes can be determined in a variety of ways, as mentioned herein. The ratio of counts on the two haplotypes can be used to determine whether there is a statistically significant difference. This ratio of counts is also referred to herein as the odds ratio. In addition, the difference between the two values can also be used; the difference can be normalized to the total number of fragments. The odds ratio and difference (and functions thereof) are examples of parameters that are compared to a threshold to determine whether there is a classification of aberration.

[0101] The RHDO analysis can utilize all alleles on the same haplotype (e.g., cumulative counts) to determine if there is any imbalance of the two haplotypes present in the plasma, e.g., can be performed in the mother's plasma, as described in Lo's patent applications 12 / 940,992 and 12 / 940,993, see above. This method can significantly increase the number of DNA molecules used to determine if there is any imbalance, can be used to distinguish cancer-caused imbalances from non-cancer stochastic fluctuations (since the allele counts in the absence of cancer or pre-malignant states are randomly distributed, cancer-caused imbalances are deterministic) and thus get better statistical power. In contrast to analyzing multiple SNP loci individually, the RHDO method can utilize the relative positions of the alleles on the two chromosomes (haplotype information), so that alleles on the same chromosome can be analyzed together. In the absence of haplotype information, the allele counts at different SNP loci cannot be added together in this statistical way to determine if a haplotype is overrepresented or underrepresented in the plasma. The quantification of allele counts can be by, but not limited to, massively parallel sequencing (e.g., Illumina's HiSeq sequencing system, Life Technologies' sequencing by ligation (SOLiD) system, Ion Torrent and Life Technologies' Ion Torrent sequencing system, nanopore sequencing (nanoporetech.com) technology, 454 sequencing technology (Roche), digital PCR (e.g., by microfluidic digital PCR (e.g., fluidigm.com)), BEAMing (bead, emulsion PCR, amplification, magnetic (inostics.com)), or droplet digital PCR (e.g., by QuantaLife (quantalife.com) and RainDance (raindancetechnologies.com)) as well as real-time PCR. In other embodiments of the technology, high-throughput target capture sequencing can be used (which utilizes liquid phase capture (e.g., utilizing Agilent SureSelect system, the Illumina TruSeq Custom Enrichment Kit (illumina.com / applications / sequencing / targeted_resequencing.ilmn), or by MyGenostics GenCap Custom Enrichment system (mygenostics.com / )) or array-based capture (e.g., using Roche NimbleGen system).

[0102] InFigure 4 In the example shown, a slight allelic imbalance is observed for the first two SNP loci (24 and 26 for the first SNP, and 18 and 20 for the second SNP). However, the number of allele counts is not statistically sufficient to determine whether a true allelic imbalance exists. Therefore, the counts of alleles on the same haplotype are added together until the cumulative counts of alleles for both haplotypes are sufficient to reach a statistical conclusion that no allelic imbalance exists between the two haplotypes for a particular subregion of the chromosomal region 410 (for this example, the fifth SNP). After reaching a statistically significant classification, the cumulative counts are reset (for this example, at the sixth SNP). The cumulative counts are then determined until the cumulative counts of alleles for both haplotypes are again sufficient to reach a statistical conclusion that no allelic imbalance exists between the two haplotypes for the particular subregion of the region 410. The total cumulative counts can also be used for the entire region, but the previous method can allow different subregions to be tested, providing a higher degree of precision in determining the location of the aberration (i.e., a subregion) than the entire region 410. Examples of statistical tests used to determine whether a true allelic imbalance exists include, but are not limited to, sequential probability ratio test (Zhou W, et al. Nat Biotechnol 2001; 19: 78-81; Zhou W, et al. Lancet. 2002; 359: 219-25), t-test, and chi-square test.

[0103] Detecting deletions

[0104] Figure 5 It is illustrated that the deletion of the chromosomal region 510 in a cancer cell and the measurement performed in the plasma, according to an embodiment of the application, allows determining the presence of the deleted region. Figure 5 The two haplotypes (Hap I and Hap II) of the chromosomal region 510 are shown for a normal cell. The haplotypes include a first plurality of heterozygous loci 520 that cover the chromosomal region 510 to be analyzed. In addition, the two haplotypes of a cancer cell are shown. In the cancer cell, the region 510 is deleted for Hap II. As Figure 4 Similarly, Figure 5 The number of allele counts for each locus 520 is also shown. In addition, the total cumulative amounts are recorded for certain subregions within the chromosomal region 510.

[0105] Since tumor tissue typically comprises a mixture of tumor and non-tumor cells, LOH can be confirmed by a relative shift in the proportion of the amount of the two alleles at loci within region 510. In this case, the deleted haplotype Hap II in region 510 can be determined by a combination of loci 520 that show a relative decrease in the amount of DNA fragments compared to the corresponding loci on normal tissue. The haplotype fragment that appears more frequently is Hap I, which is retained in the tumor cells. In certain embodiments, it is desirable to perform a process that enriches the proportion of tumor cells in the tumor sample, allowing the deleted and retained haplotypes to be more easily determined. One example of such a process is microdissection (either manually or by laser capture technology).

[0106] In theory, for a chromosomal region in tumor tissue that exhibits LOH, each allele on Hap I is present in a relative excess in peripheral blood DNA, and the degree of allelic imbalance depends on the relative percentage concentration of tumor DNA in the plasma. However, the relative amounts of the two alleles in a sample of peripheral blood DNA follow a Poisson distribution. Statistical analysis can be used to determine whether the observed allelic imbalance is due to the true presence of LOH in the cancer tissue or due to chance fluctuations. The ability to detect allelic imbalance associated with LOH in a cancer depends on the number of molecules of peripheral blood DNA analyzed and the relative percentage concentration of tumor DNA. The higher the relative percentage concentration of tumor DNA and the greater the number of molecules analyzed, the higher the sensitivity and specificity will be for detecting allelic imbalance.

[0107] In Figure 5 In the example shown, a slight allelic imbalance is observed for the first two SNP loci (24 vs. 22 for the first SNP and 18 vs. 15 for the second SNP). However, the number of allele counts is not statistically sufficient to determine whether there is a true allelic imbalance. Therefore, the counts of alleles on the same haplotype are added together until the cumulative allele counts for the two haplotypes are sufficient to make a statistical conclusion that there is an allelic imbalance between the two haplotypes on region 510 (for this example, the fifth SNP). In some embodiments, only the imbalance is known and the specific type (deletion or amplification) is not determined. The cumulative counts are then determined until the cumulative allele counts for the two haplotypes are again sufficient to make a statistical conclusion that there is an allelic imbalance between the two haplotypes for a particular subregion of region 510. The total cumulative counts can also be used for the entire region, which can be performed in any of the methods described herein.

[0108] Detection of amplification of a chromosomal region

[0109] Figure 6 It is set forth that according to embodiments of the application, the amplification of chromosomal region 610 in cancer cells and the measurement performed in plasma to determine the amplified region. In addition to LOH, amplification of chromosomal regions is frequently observed in cancer tissues. In Figure 6 In the example shown, Hap II in chromosomal region 610 is amplified to 3 copies in cancer cells. As shown, region 610 includes only 6 heterozygous loci, unlike the longer regions shown in previous figures. Of the 6 loci, amplification is identified as statistically significant, where overrepresentation is determined to be statistically significant. In some embodiments, only imbalance is known, and the specific type (deletion or amplification) is not determined. In other embodiments, cancer cells can be obtained and analyzed. Such analysis can provide information as to whether the imbalance is due to a deletion (cancer cells are homozygous for the deleted region) or amplification (cancer cells are heterozygous for the amplified region). In other embodiments, the methods described in Section IV can be used to determine whether a deletion or amplification is present, analyzing the entire region (i.e., not analyzing individual haplotypes). If the region is overrepresented, the aberration is amplification; if the region is underrepresented, the aberration is a deletion. In addition, region 620 is also analyzed, and the cumulative count demonstrates that there is no imbalance.

[0110] SPRT analysis for plasma RHDO analysis

[0111] For any chromosomal region with heterozygous loci, RHDO analysis can be used to determine whether any dosage imbalance between the two haplotypes exists in the plasma. In these regions, the presence of a haplotype dosage imbalance in the plasma indicates the presence of tumor-derived DNA in the plasma sample. In one embodiment, SPRT analysis can be used to determine whether the difference in the number of sequencing reads for Hap I and Hap II is statistically significant. In an example of SPRT analysis, we first determine the number of sequencing reads corresponding to each haplotype. Then, we determine a parameter (e.g., the percent concentration) that represents the proportion of sequencing reads contributed by the haplotype that potentially presents a relative excess (e.g., the proportion of the number of reads for one haplotype divided by the number of reads for the other haplotype). In the scenario of LOH, the haplotype that potentially presents an excess is the non-deleted haplotype, while in the scenario of monoallelic amplification of a chromosomal region, the haplotype that potentially presents an excess is the amplified haplotype. Next, this proportion is compared to two thresholds (an upper threshold and a lower threshold) that are constructed based on the null hypothesis (i.e., no haplotype dosage imbalance exists) and the alternative hypothesis (i.e., a haplotype dosage imbalance exists). If the proportion is greater than the upper threshold, it indicates that there is a statistically significant imbalance between the two haplotypes in the plasma. If the proportion is lower than the lower threshold, it indicates that there is no statistically significant imbalance between the two haplotypes. If the proportion is between the upper threshold and the lower threshold, it indicates that there is not enough statistical evidence to make a conclusion. For the region being analyzed, the number of heterozygous loci can be sequentially accumulated until SPRT classification can be successfully performed.

[0112] The mathematical formula for calculating the upper and lower thresholds for SPRT is: upper threshold = [(ln 8) / N - ln δ] / ln γ; lower threshold = [(ln 1 / 8) / N - ln δ] / ln γ, where δ = (1 - θ1) / (1 - θ2) and θ1 is the expected ratio of sequencing tags corresponding to the haplotype that potentially presents an excess when there is allelic imbalance in the plasma, θ2 is the expected ratio of either haplotype when there is no allelic imbalance, i.e., 0.5; N is the total number of sequencing tags for Hap I and Hap II; ln is the mathematical symbol that represents the natural logarithm, i.e., log e θ1 depends on the relative percent concentration (F) of tumor DNA in the plasma.

[0113] In the scenario of LOH, θ1 = 1 / (2 - F). In the scenario of monoallelic amplification, θ1 = (1 + zF) / (2 + zF), where z represents the additional increased copy number corresponding to the chromosomal region that is amplified in the tumor. For example, if one chromosome is duplicated, then it is one additional increased copy for that particular chromosome. Then z equals 1.

[0114] Figure 7 RHDO analysis of plasma DNA from HCC patients for a chromosomal segment located on chromosome lp is shown according to an embodiment of the application, where the chromosomal segment has a monoallelic amplification in tumor tissue. Green triangles represent patient data. The total number of sequencing reads increases as the number of SNPs analyzed increases. The fraction of total sequencing reads that come from the haplotype corresponding to the chromosomal amplification aberration in the tumor changes as the total number of sequencing reads analyzed increases and eventually reaches a value above the upper threshold. This indicates a significant haplotype dosage imbalance and thus corroborates the presence of this cancer-related chromosomal aberration in the plasma.

[0115] SPRT-based RHDO analysis is performed for all chromosomal regions in HCC patients where the regions show amplifications and deletions in tumor tissue. RHDO analysis is performed for 922 chromosomal segments known to have LOH and 105 chromosomal segments known to have amplification, respectively, with the following results. For LOH, 922 chromosomal segments are classified using SPRT, and 921 of them are correctly identified as having a haplotype dosage imbalance in plasma, providing 99.99% accuracy. For monoallelic amplification, 105 segments are classified using SPRT, and 105 of them are correctly identified as having a haplotype dosage imbalance in plasma, providing 100% accuracy.

[0116] C. Analysis of Relative Haplotype Size

[0117] The length of the respective segment of a haplotype can be used as an alternative to counting the dosage of segments corresponding to two haplotypes. For example, for a particular chromosomal region, the size of the DNA segment from one haplotype can be compared to the size of the DNA segment from the other haplotype. One can analyze the size distribution of the DNA segment corresponding to either allele at the heterozygous locus of the first haplotype of the region and compare it to the size distribution of the DNA segment corresponding to either allele at the heterozygous locus of the second haplotype. A statistically significant difference in the size distribution can be used to identify an aberration in the same way as dosage counting.

[0118] It has been reported that the size distribution of total plasma DNA (i.e., tumor vs. non-tumor DNA) is elevated in cancer patients (Wang BG, et al. Cancer Res. 2003; 63:3966-8). However, when tumor-derived DNA (not total DNA, i.e., tumor vs. non-tumor DNA) is specifically studied, the length distribution of tumor-derived DNA molecules is observed to be smaller than that of molecules derived from non-tumor cells (Diehl et al. Proc Natl Acad Sci US A. 2005; 102:16368-73). Therefore, the size distribution of peripheral blood DNA can be used to determine the presence of cancer-related chromosomal aberrations. The principle of size analysis is shown in... Figure 8 middle.

[0119] Figure 8 The illustration shows a variation in fragment size distribution across two haplotypes of a chromosomal region when a tumor including deletion aberrations is present, according to an embodiment of the invention. Figure 8 As shown, the T allele is deleted in tumor tissue. Consequently, tumor tissue releases only short molecules of the A allele into the plasma. The tumor-derived short DNA molecules cause an overall shortening of the length distribution corresponding to the A allele in the plasma, thus making the length distribution of the A allele in plasma shorter than that of the T allele. As discussed in previous sections, all alleles located on the same haplotype can be analyzed together. In other words, the size distribution of DNA molecules carrying alleles on one haplotype can be compared with the size distribution of DNA molecules carrying alleles on other haplotypes. The haplotype missing in tumor tissue shows a longer size distribution in plasma.

[0120] In addition, size analysis can also be used to detect amplification of chromosomal regions associated with cancer. Figure 9 This illustration shows a variation in fragment size distribution across two haplotypes of a chromosomal region in the presence of a tumor including amplified aberrations, according to an embodiment of the invention. Figure 9 In the example shown, the chromosomal region carrying the T allele is replicated in the tumor. As a result, an increased amount of the shorter DNA molecule carrying the T allele is released into the plasma, thus making the size distribution of the T allele-corresponding fragment appear shorter overall than that of the A allele-corresponding fragment. Similarly, all alleles located on the same haplotype can be analyzed together. In other words, the size distribution of haplotypes amplified in tumor tissue appears shorter than the size distribution of haplotypes not amplified in the tumor.

[0121] Detection of shortened size distribution of peripheral blood DNA

[0122] The size of the DNA fragments obtained from the two haplotypes (i.e. Hap I and Hap II) can be determined by, but not limited to, paired-end massively parallel sequencing. After sequencing the ends of the DNA fragments, the sequencing reads (tags) can be aligned to the human reference genome. The size of the sequenced DNA molecules can be deduced from the coordinates of the outermost nucleotides at each end. The sequencing tags of the molecules can be used to determine whether the sequenced DNA fragments were obtained from Hap I or Hap II. For example, one of the sequencing tags can include a heterozygous locus in the chromosomal region to be analyzed.

[0123] Thus, for each sequenced molecule, we can determine its length and whether it was obtained from Hap I or Hap II. Based on the size of the fragments aligned to the respective haplotypes, a computer system can calculate the size distribution (e.g. average fragment size) of Hap I and Hap II. Suitable statistical analyses can be used to compare the size distribution of the DNA fragments obtained from Hap I and Hap II to determine whether the size distribution is different enough to allow discrimination of the presence or absence of aberration. In addition to paired-end massively parallel sequencing, other methods can be used to determine the size of the DNA fragments, including but not limited to sequencing the entire DNA fragments, mass spectrometry, and optical methods for observing and comparing the length of the observed DNA molecules to a standard.

[0124] Next, we introduce two example methods for determining the shortening of the peripheral blood DNA associated with genetic aberrations of tumors. The purpose of the two methods is to provide a quantitative measure of the difference in size distribution of the DNA fragments of the two populations. The DNA fragments of the two populations refer to the DNA molecules corresponding to Hap I and Hap II.

[0125] Difference in the proportion of short DNA fragments

[0126] In one embodiment, the proportion of short DNA fragments is used. A length threshold (w) is set to define short DNA molecules. Different length thresholds can be varied and selected to fit different diagnostic purposes. The computer system can determine the number of molecules that are equal to or shorter than the length threshold. Then, the proportion of short DNA fragments (Q) can be calculated by dividing the number of short DNA by the total number of DNA fragments. The Q value is affected by the length distribution of the population of DNA molecules. A shorter overall size distribution indicates a higher proportion of DNA molecules as short fragments, thus a higher Q value.

[0127] Next, the difference in the corresponding proportion of short DNA fragments between Hap I and Hap II can be used. The difference in the length distribution of DNA fragments obtained by Hap I and Hap II (ΔQ) can be reflected by the difference in the proportion of short fragments for Hap I and Hap II. ΔQ = Q HapI - Q HapII where Q HapI is the proportion of short fragments corresponding to Hap I DNA fragments, and Q HapII is the proportion of short fragments corresponding to Hap II DNA fragments. Q HapI and Q HapII are examples of statistical values of the length distribution of fragments corresponding to each of the two haplotypes.

[0128] As shown in the previous section, when Hap II is deleted in tumor tissue, the length distribution for Hap I DNA fragments is shorter than the length distribution for Hap II DNA fragments. As a result, a positive ΔQ value is observed. The positive ΔQ value can be compared to a threshold value, to determine whether ΔQ is large enough to conclude that a deletion is present. Amplification of Hap I also shows a positive ΔQ value. When there is duplication of Hap II in tumor tissue, the length distribution for Hap II DNA fragments is smaller than the length distribution for Hap I DNA fragments. Therefore, the ΔQ value is called negative. In the absence of chromosomal aberrations, the length distribution for Hap I and Hap II DNA fragments in plasma / serum is similar. Therefore, the ΔQ value is approximately zero.

[0129] The ΔQ of a patient can be compared to that of a normal individual, to determine whether the value is normal. In addition, the ΔQ value of a patient can be compared to that obtained from a patient with a similar cancer, to determine whether the value is abnormal. Such comparisons can involve comparison to a threshold value as described herein. In terms of disease monitoring, the value ΔQ can be monitored continuously over a period of time. Changes in the ΔQ value can indicate an increase or decrease in the relative percentage concentration of tumor DNA in plasma / serum. In selected embodiments of the technology, the relative percentage concentration of tumor DNA can be related to the stage of the tumor, prognosis and progression of the disease. The use of such embodiments in measuring methods at different time points will be discussed in more detail below.

[0130] Difference in the proportion of total length contributed by short DNA fragments

[0131] In this embodiment, the proportion of the total length contributed by short DNA fragments is used. The computer system can determine the total length of a set of DNA fragments in the sample (e.g. from a particular haplotype of a given region or from fragments obtained from a given region only). Different cutoff sizes (w) can be chosen, below which DNA fragments are defined as "short fragments". Different cutoff sizes can be varied and chosen to fit different diagnostic purposes. Then, the computer system can determine the total length of short DNA fragments by adding up the lengths of randomly selected DNA fragments that are equal to or less than the cutoff size. The proportion of the total length contributed by short DNA fragments is then calculated as follows: F =∑ w length / ∑ N length, where∑ w lengthrepresents the sum of the lengths of DNA fragments that are equal to or less than w (bp); and∑ N lengthrepresents the sum of the lengths of DNA fragments that are equal to or less than a predetermined length N. In one embodiment, N is 600 bases. However, other length definitions, such as 150 bases, 180 bases, 200 bases, 250 bases, 300 bases, 400 bases, 500 bases, and 700 bases, can be used to calculate the "total length".

[0132] Since the Illumina Genome Analyzer system is not very efficient in amplifying and sequencing DNA fragments longer than 600 bases, N can be chosen to be the value of 600 bases. In addition, limiting the analysis to DNA fragments less than 600 bases can also avoid bias in the measurement due to genomic structural variations. In the presence of structural variations, such as rearrangements (Kidd JM et al, Nature 2008; 453: 56-64), the length of a DNA fragment can be overestimated when the end of the DNA fragment is aligned to the reference genome by bioinformatics means. In addition, among all DNA fragments that successfully align to the reference genome, more than 99.9% of the fragments are less than 600 bases, therefore, including all fragments that are equal to and less than 600 bases will provide an unbiased estimate of the length distribution of DNA fragments in the sample.

[0133] Therefore, the difference in the proportion of the total length contributed by short DNA fragments between Hap I and Hap II can be used. The change in the length distribution between Hap I and Hap II DNA fragments can be reflected by their F values. Here, we will refer to F Hap I and F Hap IIFhap I and Fhap II are defined as the fraction of short DNA fragments corresponding to Hap I and Hap II, respectively. The difference in the fraction of short DNA fragments between Hap I and Hap II (ΔF) can be calculated as follows: ΔF = Fhap I - Fhap II Hap I - F Hap II . F Hap I and F Hap II are statistical values of the two sets of fragment length distributions corresponding to each haplotype.

[0134] Similar to the embodiments shown in the above section, deletion of Hap II in tumor tissue results in an apparent shortening of the Hap I DNA length distribution when compared to Hap II DNA fragments. This results in a positive ΔF value. When Hap II is duplicated, a negative ΔF value is observed. In the absence of chromosomal aberrations, the ΔF value is close to zero.

[0135] The patient's ΔF can be compared to that of a normal individual to determine if the value is normal. The patient's ΔF can be compared to that obtained from a patient with a similar cancer to determine if the value is abnormal. Such comparisons can involve comparison to a threshold value as described herein. In terms of disease monitoring, the ΔQ value can be monitored serially. Changes in the ΔF value can indicate an increase or decrease in the percentage concentration of tumor DNA in the plasma / serum.

[0136] D. General Methods

[0137] Figure 10 Figure 1 is a flow chart showing a method for analyzing a biological sample of an organism for haplotypes to determine if a chromosomal region exhibits a deletion or amplification according to an embodiment of the present application. The biological sample includes nucleic acid molecules (also referred to as fragments) derived from normal cells as well as potentially cancer-related cells. These molecules can be free in the sample. The organism can be any type having more than one copy of a chromosome, i.e., at least a diploid organism, but can include higher polyploid organisms.

[0138] In one embodiment of this and any other method described herein, the biological sample comprises free DNA fragments. While the analysis of plasma DNA is used to illustrate the different methods described in this application, these methods can also be applied to detect tumor-related chromosomal aberrations in samples comprising a mixture of normal and tumor-derived DNA. Other sample types include saliva, tears, pleural fluid, ascites, bile, urine, serum, pancreatic juice, stool, and cervical smear samples.

[0139] In step 1010, first and second haplotypes are determined for normal cells of the organism at a first chromosomal region. Haplotypes can be determined by any suitable method, such as those mentioned herein. The chromosomal region can be selected by any method, such as those described herein. The first chromosomal region includes a first plurality of loci (e.g., loci 420 in region 410), and it is heterozygous. The heterozygous loci (hets) can be separated from each other by a distance, e.g., in the first plurality of loci, a locus can be 500 or 1000 bases (or more) from another locus. Other heterozygous loci (hets) can be present in the first chromosomal region, but are not necessarily utilized.

[0140] In step 1020, each of the plurality of nucleic acid molecules in the biological sample is characterized by a location and an allele. For example, the nucleic acid molecule can be identified in a reference genome of the organism. This localization can be achieved in a variety of ways, including using molecular sequencing (e.g., by universal sequencing) to obtain one or two (paired-end) sequencing tags of the molecule, and then aligning the sequencing tags to a reference genome. Such alignment can be performed using tools such as Basic Local Alignment Search Tool (BLAST). This localization can be identified as a numerical value in a chromosome arm. The allele at a heterozygous locus (het) can be used to identify which haplotype the fragment is from.

[0141] In step 1030, a first set of nucleic acid molecules is identified as being from the first haplotype based on the identified locations and determined alleles. For example, the first set of nucleic acid molecules includes Figure 4 The fragment of locus 421 (which has allele A) shown is identified as being from Hap I. The first set of nucleic acid molecules can cover the first chromosomal region as long as it includes at least one nucleic acid molecule that is localized at each of the first plurality of loci.

[0142] In step 1040, a second set of nucleic acid molecules is identified as being from the second haplotype based on the identified locations and determined alleles. For example, the second set of nucleic acid molecules includes Figure 4 The fragment of locus 421 (which has allele T) shown is identified as being from Hap II. The second set includes at least one nucleic acid molecule that is localized at each of the first plurality of loci.

[0143] In step 1050, the computer system calculates a first value for the first set of nucleic acid molecules. The first value defines a property of the first set of nucleic acid molecules. Examples of first values include a tag count corresponding to the number of nucleic acid molecules in the first set, and a length distribution corresponding to the first set of nucleic acid molecules.

[0144] In step 1060, the computer system calculates a second value for the second set of nucleic acid molecules. The second value defines a property of the second set of nucleic acid molecules.

[0145] In step 1070, the first value is compared to the second value to determine a classification of whether the first chromosomal region exhibits a deletion or amplification. The classification of the presence of a deletion or amplification can provide information that the organism has cells associated with cancer. Examples of comparisons include considering the difference or ratio of the two values, and comparing the result to one or more thresholds, as described herein. For example, in an SPRT analysis, the ratio can be compared to a threshold. Example classifications can include positive (i.e., an amplification or deletion is detected), negative, undetermined, and a degree of change corresponding to positive and negative (e.g., using an integer from 1 to 10, or a real number from 0 to 1). Amplification can include simple duplication. Such methods can detect the presence of cancer-associated nucleic acids, including DNA from tumor DNA and pre-neoplastic lesions (i.e., precursors to cancer).

[0146] E. Depth

[0147] The depth of analysis refers to the number of molecules that need to be analyzed to provide a classification or other determination to a specified degree of accuracy. In one embodiment, the depth can be calculated based on known aberrations, and then measurements and analysis are performed under conditions that satisfy the depth. In another embodiment, analysis can be performed continuously until a successful classification, and the depth corresponding to the successful classification can be used to determine a level of cancer (e.g., a stage of cancer or size of a tumor). Examples of some calculations related to depth are provided below.

[0148] As described herein, a bias can refer to any difference or ratio. For example, a bias can be between a first value and a second value, or a first parameter and a second parameter, resulting from a threshold or tumor concentration. If the bias is doubled, the number of fragments that need to be measured is reduced to 1 / 4. More generally, if the bias is increased to N times, the number of fragments that need to be measured is 1 / N of the original. 2 From this, it follows that if the bias is reduced to 1 / N, the number of fragments to be examined is increased to N times the original. 2 N can be a real number or an integer.

[0149] Consider a case where the tumor DNA is 10% of the sample (e.g., plasma), and assume that 10 million fragments are sequenced and a statistically significant difference is observed. Then, for example, an enrichment process is performed so that the sample has 20% tumor DNA, and then 2.5 million fragments are needed. In this way, the depth can be related to the percentage concentration of tumor DNA in the sample.

[0150] Furthermore, the amount of amplification also affects the depth. For a region that is amplified to twice the original normal copy number (e.g., 4, as opposed to 2, which is normal), assume that X number of fragments are needed for analysis. If the region is amplified to 4 times the original, then X / 4 number of fragments are needed for the region.

[0151] F. Threshold

[0152] As described above, the deviation or bias of a parameter from a normal value (e.g., the difference value or ratio between haplotypes) can be used to provide a diagnosis. For example, the bias can be the difference between the average size of fragments from one haplotype of a region and the average size of fragments from other haplotypes. If the bias is greater than a certain amount (e.g., a threshold value determined from normal samples and / or regions), then a deletion or amplification is identified. However, the degree to which the threshold is exceeded can be useful, and thus multiple thresholds can be used, each corresponding to a different level of cancer. For example, a higher bias from a normal value can provide an indication of the stage of cancer (e.g., a higher degree of imbalance for stage 4 than for stage 3). Furthermore, a higher bias can also be due to a larger tumor releasing many fragments, and / or the region analyzed being amplified multiple times.

[0153] In addition to providing different levels of cancer, varying the threshold allows for efficient detection of regions or specific regions that have aberrations. For example, one can set a high threshold, and thus look primarily for amplifications of 3 times or more, which would result in a higher imbalance than a deletion of one haplotype. In addition, one can also detect deletions of regions that have 2 copies. Furthermore, a lower threshold can be used to identify regions that can have aberrations, and then further analysis can be performed on these regions to determine whether an aberration exists and where it exists. For example, a binary search (or higher level search, such as octree search) can be used, with a higher threshold used at a lower level in the search hierarchy.

[0154] Figure 11 A region 1110 in a cancer cell with a deletion aberration in a subregion 1130 is shown, as well as a measurement in plasma to determine the deleted region according to an embodiment of the application. The chromosomal region 1110 can be selected by any of the methods mentioned herein, such as by splitting the genome into equal sized bins. In addition, for each of the loci 1120, Figure 11 The number of allele counts is also shown. In addition, the cumulative total is also recorded for region 1140 (normal region) and region 1130 (deleted region), respectively.

[0155] If region 1110 is selected for analysis, the cumulative count of Hap I is 258, while the cumulative count of Hap II is 240, providing a difference of 18 over 11 loci. The percentage of this difference over the total number of counts is much smaller than if only the corresponding sub-region 1130 of the deletion is analyzed. This is reasonable because about half of the regions in region 1110 are normal while all of sub-region 1130 is deleted in the cancer cell. Thus, aberrations in region 1110 can be missed depending on the threshold used.

[0156] To allow detection of deletions of sub-regions, for relatively large regions (in this example, assume that region 1110 is relatively larger than the size of the deletion region to be identified), multiple embodiments can use a lower threshold. A lower threshold will identify more regions, which can include some false positives, but will reduce false negatives. Currently, false positives can be removed by further analysis, which can also pinpoint the aberration.

[0157] Once a region is identified for further analysis, the region can be divided into different sub-regions for further analysis. Figure 11 In this example, one can divide the 11 loci into two halves (e.g., using a binary tree), providing a sub-region 1140 of 6 loci and a sub-region 1130 of 5 loci. The same threshold or a more stringent threshold can be used to analyze these regions. In this example, sub-region 1140 is then identified as normal, while sub-region 1130 is identified as including a deletion or amplification. In this way, a larger region can be dismissed from having an aberration and time can be spent to further analyze regions that are suspicious (regions above the lower threshold) to identify sub-regions that show aberrations with high reliability (e.g., using a higher threshold). Although RHDO is used herein, the size technique is equally applicable.

[0158] The size of the regions (and the size of the sub-regions at lower levels in the tree) used for the first level of search can be selected based on the size of the aberrations to be detected. Studies have found that cancers show 10 aberrant regions with lengths of 10 MB. In addition, patients have 100 MB of aberrant regions. Later stage cancers can have larger size aberrations.

[0159] G. Refinement of aberration location within a region

[0160] In the previous section, the division of regions into sub-regions based on tree search was discussed. Here, other methods for analyzing sub-regions and how to pinpoint aberrations within a region are discussed.

[0161] Figure 12It is shown how the location of the distortion can be mapped using RHDO analysis according to embodiments of the application. The horizontal shows the chromosome region, where the haplotypes of the non-cancerous cells are labeled Hap I and Hap II. The region of loss of Hap II in the cancerous cells is labeled LOH.

[0162] As shown, the RHDO analysis starts on the left side of the hypothetical chromosome region 1202 and proceeds to the right side. Each arrow represents a classified chromosome segment of RHDO. Each chromosome segment can be considered as its own region, specifically a sub-region of the larger heterozygote region. Prior to classification determination, the size of the RDDO classified segment depends on the number of loci (and the location of the loci). The number of loci included in each RHDO segment depends on the number of molecules used for analysis of the respective segment, the required reliability (e.g. odds ratio in SPRT), and the relative percentage concentration of tumor-derived DNA in the sample. As Figure 4 and Figure 5 shown in the example, when the number of molecules is sufficient to determine that there is a statistically significant difference between the two haplotypes, then a classification is made.

[0163] Each solid horizontal arrow represents a RHDO classified segment, which shows that there is no imbalance in haplotype dosage in the DNA sample. In the region of the tumor that does not have LOH, 6 RHDO classifications are made, and each classification represents that there is no imbalance in haplotype dosage. The next RHDO classified segment 1210 crosses the junction 1205 between the region with LOH and the region without LOH. In the upper portion of Figure 12 , the SPRT curve for RHDO segment 1210 is shown. The black vertical arrow represents the junction between the region with LOH and the region without LOH. As more and more data accumulates from the region with LOH, the RHDO classification of this chromosome segment indicates that there is an imbalance in haplotype dosage.

[0164] Each white horizontal arrow represents a RHDO classified segment, which indicates that there is an imbalance in haplotype dosage. Furthermore, the next 4 RHDO on the right indicate that there is an imbalance in haplotype dosage in the DNA sample. It can be inferred that the junction between the region with LOH and the region without LOH is located in the first segment that shows a change in RHDO classification, i.e. from an imbalance in haplotype dosage to a lack of imbalance in haplotype dosage, or vice versa.

[0165] Figure 13 It is shown how the location of the distortion can be mapped using RHDO analysis according to embodiments of the application. The horizontal shows the chromosome region, where the haplotypes of the non-cancerous cells are labeled Hap I and Hap II. The region of loss of Hap II in the cancerous cells is labeled LOH. Figure 13In this case, the RHDO classification from both directions is shown. With the RHDO analysis starting from the left, it can be inferred that the junction between the region with LOH and the region without LOH lies within the first RHDO segment 1310 that shows a haplotype dosage imbalance. With the RHDO analysis starting from the right, it can be inferred that the junction lies within the first RHDO segment 1320 that shows a lack of haplotype dosage imbalance. Combining the information from the RHDO analysis performed from both directions, it can be inferred that the junction between the region with LOH and the region without LOH lies at position 1330.

[0166] IV. Detection of aberrant non-specific haplotypes

[0167] The RHDO method relies on the use of heterozygous loci. Now, chromosomes of diploid organisms have some variation, resulting in two haplotypes, but the number of heterozygous loci is variable. Some individuals can have relatively few heterozygous loci. In addition, the embodiments described in this section can also be used for homozygous loci, where two haplotypes on the same region are compared by comparing two regions. Thus, although there can be some disadvantages due to two different chromosomal regions, more data points can be obtained.

[0168] In the relative chromosomal region dosage method, the number of fragments obtained from a chromosomal region (e.g., determined by counting sequencing tags that align to the region) is compared to an expected value (which can be obtained from a reference chromosomal region or from the same region in another sample known to be healthy). In this way, fragments can be counted for a chromosomal region regardless of which haplotype the sequencing tag came from. Thus, sequencing tags that do not contain a heterozygous site can still be used. To perform the comparison, the embodiments can normalize the tag counts prior to comparison. Each region is defined by at least two loci (the two loci are separated from each other), and fragments at these loci can be used to obtain an overall value for the region.

[0169] For a particular region, the normalized value of the sequencing reads (tags) can be calculated by dividing the number of sequencing reads aligned to the region by the number of sequencing reads aligned to the entire genome. This normalized tag count allows the results from one sample to be compared to the results from another sample. For example, the normalized value can be the expected proportion (e.g., percentage or fraction) of sequencing reads from a particular region, as described above. However, many other normalizations are possible, as will be apparent to one skilled in the art. For example, one can normalize by dividing the number of counts for a region by the number of counts for a reference region (in the above case, the reference region is the entire genome). The normalized tag count can then be compared to a threshold value, which can be determined from one or more reference samples that do not have cancer.

[0170] Next, the normalized tag count of the test individual is compared to the normalized tag count of one or more reference subjects (e.g., that do not have cancer). In one embodiment, the comparison is made by calculating the z-score of the test individual for a particular chromosomal region. The z-score is calculated using the following equation: z-score = (normalized tag count in the case - mean) / S.D., where "mean" is the average of the normalized tag counts aligned to the particular chromosomal region for the reference samples; and S.D. is the standard deviation of the normalized tag counts aligned to the particular region for the reference samples. Thus, the z-score is the number of standard deviations that the normalized tag count of the test individual is from the average of the normalized tag counts of the reference subject(s) for a particular chromosomal region.

[0171] In the case where the organism being tested has cancer, chromosomal regions that are amplified in the tumor tissue are over-represented in the plasma DNA. This results in a positive z-score. On the other hand, chromosomal regions that are deleted in the tumor tissue are under-represented in the plasma DNA. This results in a negative z-score. The magnitude of the z-score depends on a number of factors.

[0172] One factor is the relative percentage concentration of tumor-derived DNA in the biological sample (e.g., plasma). The higher the relative percentage concentration of tumor-derived DNA in the sample (e.g., plasma), the greater the difference between the normalized tag count of the test individual and the count of the reference individual. Thus, a greater magnitude z-score is obtained.

[0173] Another factor is the degree of fluctuation in the normalization tag counts in the one or more reference individuals. In the biological sample (e.g., plasma) of the subject being tested, the same degree of overpresentation of a chromosomal region will result in a larger z-score with less fluctuation (i.e., smaller standard deviation) in the normalization tag counts in the reference set. Similarly, in the biological sample (e.g., plasma) of the subject being tested, the same degree of underpresentation of a chromosomal region will result in a larger absolute value negative z-score with less fluctuation (i.e., smaller standard deviation) in the normalization tag counts in the reference set.

[0174] Another factor is the degree of chromosomal aberration in the tumor tissue. For a particular chromosomal region, the degree of chromosomal aberration refers to the change in copy number (gain or loss). The greater the change in copy number in the tumor tissue, the greater the degree of overpresentation or underpresentation of the particular chromosomal region in the plasma DNA. For example, the loss of two copies of a chromosome will result in a greater degree of underpresentation of the chromosomal region in the plasma than the loss of one of two copies of the chromosome, and thus a larger absolute value negative z-score. Typically, there are multiple chromosomal aberrations in a cancer. Chromosomal aberrations can further vary among various cancers in their nature (i.e., amplification or deletion), their degree (single or multiple copy gain or loss), and their length (size of the aberration).

[0175] The precision of measuring the normalization tag counts can be affected by the number of molecules analyzed. We estimate that 15,000, 60,000, and 240,000 molecules need to be analyzed to detect a chromosomal aberration (gain or loss) with one copy change at percent concentrations of approximately 12.5%, 6.3%, and 3.2%, respectively. Details of the tag counts used to detect cancer for different chromosomal regions are described in U.S. Patent Publication No. 2009 / 0029377 (Lo et al.) entitled "Diagnosing Fetal Chromosomal Aneuploidy Using Massively Parallel Genomic Sequencing," the entire contents of which are incorporated herein by reference for all purposes.

[0176] In addition, multiple embodiments can also use length analysis without the use of tag counts. Furthermore, length analysis can also be used without normalized tag counts. As mentioned herein and described in U.S. Patent Application No. 12 / 940,992, length analysis can use multiple parameters. For example, the Q or F values derived above can be used. Since such length values are not proportional to the number of reads, these values need not be normalized by counts from other regions. Techniques for haplotype specific methods can also be applied to non-specific methods. For example, techniques involving the depth of a region and refinements can be used. In some embodiments, corrections based on the GC content of a particular region can be considered when comparing two regions. Since the RHDO method uses the same regions, such corrections are not needed.

[0177] V. Multiple Regions

[0178] Although certain cancers can typically be associated with a particular chromosomal region, such cancers do not always present in only the same region. For example, other chromosomal regions can show aberrations, and the location of such other regions can be unknown. Furthermore, when screening for cancer in order to identify early stage cancers, one can want to identify multiple cancers that can present aberrations anywhere in the genome. To address these situations, multiple embodiments can be used to systematically analyze multiple regions in order to determine which regions show aberrations. For example, the number of aberrations and their location (e.g., whether they are contiguous) can be used to confirm the aberration, the stage of the cancer, provide a diagnosis of the cancer (e.g., whether the number is greater than a certain threshold), and provide a prognosis based on the number and location of the regions that present aberrations.

[0179] Accordingly, multiple embodiments can identify whether an organism has cancer based on the number of regions that show aberrations. Thus, one can examine multiple regions (e.g., 3000) in order to identify the number of regions that show aberrations. The regions can cover the entire genome or only a portion of the genome, such as non-repetitive regions.

[0180] Figure 14 A flowchart of a method 1400 for analyzing a biological sample of an organism using multiple chromosomal regions in accordance with embodiments of the present application. The biological sample includes nucleic acid molecules (also referred to as fragments).

[0181] In step 1410, the computer system identifies a plurality of non-overlapping chromosomal regions of the organism. Each chromosomal region includes a plurality of loci. As mentioned above, the regions can be 1 Mb in size, or some other equivalent size. The entire genome can include approximately 3000 regions, each having a predetermined size and location. Further, as mentioned above, such predetermined regions can vary to accommodate a particular length of a particular chromosome or a particular number of regions to be used, as well as any other criteria mentioned herein. If the regions have different lengths, such lengths can be used to normalize the results, e.g., as described herein.

[0182] In step 1420, the computer system identifies, for each of the plurality of nucleic acid molecules, its location in the reference genome of the organism. The location can be determined in any manner mentioned herein, e.g., by sequencing the fragment to obtain a sequencing tag and aligning the sequencing tag to the reference genome. Further, a particular haplotype of the molecule can be detected and haplotype-specific methods can be used.

[0183] Steps 1430-1450 are performed for each of the chromosomal regions. In step 1430, based on the identified locations, the computer system identifies a set of nucleic acid molecules corresponding to the chromosomal region. The set of nucleic acid molecules includes at least one nucleic acid molecule at each locus of the chromosomal region. In one embodiment, the set can be a fragment specific to a particular haplotype of the chromosomal region, e.g., as described in the RHDO method described above. In another embodiment, the set can be any fragment specific to the chromosomal region, e.g., the methods described in Section IV.

[0184] In step 1440, the computer system calculates a value for each set of nucleic acid molecules. The value defines a property of the set of nucleic acid molecules. The value can be any value mentioned herein. For example, the value can be the number of fragments in the set or a statistical value of the length distribution of the fragments in the set. The value can also be a normalized value, e.g., the tag count of the region divided by the total number of tag counts of the sample, or divided by the number of tag counts of the reference region. Further, the value can be a difference or ratio value (e.g., in RHDO), whereby the properties of different regions are compared.

[0185] In step 1450, the values are compared to a reference value to determine a classification of whether the first chromosomal region exhibits a deletion or amplification. The reference value can be any threshold or reference value described herein. For example, the reference value can be a threshold value determined for normal samples. For RHDO, the values can be the difference or ratio of the tag counts for the two haplotypes, and the reference value can be a threshold value for determining the presence of a statistically significant difference. For example, the reference value can be the tag count or size of the other haplotype or region, and the comparison can include, but is not limited to, the difference or ratio (or a function of such values), and then determining whether the difference or ratio is greater than the threshold value.

[0186] The reference value can change based on the results of other regions. For example, if neighboring regions also show deviations (albeit smaller, e.g., a z-score of 3, compared to one threshold value), a lower threshold value can be used. For example, if 3 consecutive regions are above a first threshold value, the likelihood of cancer is greater. Thus, the first threshold value can be lower than another threshold value necessary to identify cancer from non-consecutive regions. Having 3 regions (or more than 3 regions) with small deviations can reduce the likelihood of random fluctuations to a sufficiently low level, such that a high sensitivity and specificity can be maintained.

[0187] In step 1460, the number of chromosomal regions classified as exhibiting a deletion or amplification is determined. There can be some restrictions on the counted chromosomal regions. For example, only regions that are contiguous with at least one other region can be counted (or contiguous regions can be required to reach a certain size, e.g., 4 or more regions). For embodiments in which regions are not equal, the number can take into account the respective lengths (e.g., the number can be the total length of aberrant regions).

[0188] In step 1470, the number is compared to a threshold value for the number to determine a classification of the sample. For example, the classification can be whether the organism has cancer, the stage of cancer, and the prognosis of cancer. In one embodiment, all aberrant regions are counted and a single threshold value is used, regardless of where the regions occur. In another embodiment, the threshold value can change based on the location and size of the counted regions. For example, the number of regions on a particular chromosome or chromosome arm can be compared to a threshold value for that particular chromosome (or arm). Multiple threshold values can be used. For example, the number of aberrant regions on a particular chromosome (or arm) must be greater than a first threshold value, and the total number of aberrant regions in the genome must be greater than a second threshold value.

[0189] The threshold for the number of regions can also depend on how imbalanced the counted regions are. For example, the threshold for the number of regions for determining a cancer classification can depend on the specificity and sensitivity of determining aberrations in each region (aberration threshold). For example, if the aberration threshold is low (e.g., z-score of 2), a high threshold for the number of regions can be selected (e.g., 150). However, if the aberration threshold is high (e.g., z-score of 3), the threshold for the number of regions can be lower (e.g., 50). In addition, the number of regions displaying aberrations can also be weighted, e.g., regions displaying high imbalance can be given a higher weight (i.e., there are more classifications than just aberration positive and negative) than regions displaying less imbalance.

[0190] Thus, the number of chromosomal regions (which can include the number and / or size) that exhibit a significant over- or under-representation of the corresponding normalized tag counts (or other values corresponding to the attributes of the group) can be used to reflect the severity of the disease. The number of chromosomal regions displaying aberrations from the normalized tag counts can be determined by two factors, namely, the number (or size) of chromosomal aberrations in the tumor tissue, and the percentage concentration of tumor-derived DNA in the biological sample (e.g., plasma). More advanced cancers tend to display more (and larger) chromosomal aberrations. Thus, larger chromosomal aberrations associated with cancer can potentially be detected. In patients with more advanced cancers, the higher tumor burden results in a higher percentage concentration of tumor-derived DNA in the plasma. As a result, tumor-associated chromosomal aberrations are more likely to be detected in plasma samples.

[0191] In the context of cancer screening or detection, the number of chromosomal regions (which can include the number and / or size) that exhibit a significant over- or under-representation of the corresponding normalized tag counts (or other values) can be used to determine the likelihood that the tested subject has cancer. Using a cutoff of ±2 (i.e., z-score > 2 or <-2), about 5% of the tested regions are expected to show a significant deviation from the mean of the control subjects due to chance random fluctuations. When the entire genome is divided into 1 Mb segments, there are about 3000 chromosomal segments in the entire genome. Thus, about 150 chromosomal segments are expected to have a z-score of > 2 or <-2 due to random fluctuations.

[0192] Thus, for the number of segments with z-scores > 2 or <-2, a threshold of 150 can be used to determine the presence of cancer. For the number of segments with aberrant z-scores (e.g., 100, 125, 175, 200, 250, and 300), other cutoffs can be chosen to fit the diagnostic purpose. A lower cutoff (e.g., 100) will result in a higher sensitivity test but a lower specificity, while a higher cutoff will result in a higher specificity but a lower sensitivity. The number of false positive calls can be reduced by increasing the cutoff for the z-score. For example, if the cutoff is increased to 3, then only 0.3% of the segments are false positives. In this case, the presence of cancer can be indicated by the presence of more than 3 segments with aberrant z-scores. In addition, other thresholds can be chosen, e.g., 1, 2, 4, 5, 10, 20, and 30, to fit different diagnostic purposes. However, as the number of aberrant segments required for a diagnosis increases, the sensitivity of detecting chromosomal aberrations associated with cancer decreases.

[0193] One possible approach to improve sensitivity without sacrificing specificity would be to consider the results of adjacent chromosomal segments. In one embodiment, the cutoff for the z-score remains > 2 and <-2. However, only when two consecutive segments show the same type of aberration, e.g., both segments have a z-score > 2, will the chromosomal region be classified as potentially aberrant. If the deviation of the normalized tag counts is random error, then the probability of having two consecutive segments (false positives in the same direction) is 0.125% (5% x 5% / 2). On the other hand, if the chromosomal aberration encompasses two consecutive segments, a lower cutoff will make the test more sensitive to over- or under-representation of segments in the plasma sample. Since the deviation of the normalized tag counts (or other values) from the average of the control subjects is not due to random error, the requirement of consecutive calls will not have a significant detrimental effect on sensitivity. In other embodiments, the z-scores of adjacent segments can be added together if a higher cutoff is used. For example, the z-scores of 3 consecutive segments can be summed and a cutoff of 5 can be used. This concept can be extended to more than 3 consecutive segments.

[0194] In addition, the combination of the number and aberration thresholds can also depend on the purpose of the analysis and any prior knowledge (or lack thereof) of the organism. For example, if cancer screening is performed on a normal, healthy population, then a high specificity is usually used for the number of regions (i.e., a high threshold is used for the number of regions) and the aberration threshold to identify whether a region has a chromosomal aberration. However, in patients with a higher risk (e.g., patients with a mass or family history, smoking, HPV virus, hepatitis virus, or other viruses), then the thresholds can be lowered to have a higher sensitivity (lower false negatives).

[0195] In one embodiment, if 1 Mb resolution and tumor derived DNA with a detection limit of 6.3% is used to detect chromosomal aberrations, the number of molecules per 1 Mb bin must be 60,000. For the entire genome, then, approximately 180,000,000 (60,000 reads / Mb x 3,000 Mb) aligned reads are needed.

[0196] Figure 15 An embodiment according to the present application is shown in table 1500, which illustrates different numbers of chromosomal bins and the depth required for the relative percentage concentration of tumor derived fragments. Column 1510 provides the concentration of DNA fragments derived from tumor cells of a sample. The higher the concentration, the easier it is to detect aberrations, and thus the fewer molecules needed for analysis. Column 1520 provides the estimated number of molecules needed per chromosomal bin, which can be calculated by the methods described above in the section on depth.

[0197] Smaller sized chromosomal bins result in higher resolution, which is suitable for detecting smaller chromosomal aberrations. However, this requires an increased number of molecules for analysis. Larger sized chromosomal bins reduce the number of molecules needed for analysis at the cost of resolution. Thus, only larger aberrations can be detected. In one embodiment, the larger the region used, the chromosomal bin showing the aberration will be subdivided and analyzed for these sub-regions, resulting in higher resolution (e.g. as described above). Column 1530 provides the size of the respective chromosomal bins. The smaller the value, the larger the number of corresponding regions. Column 1540 shows the number of molecules needed for the entire genome. Thus, if one has the evaluation parameters (or minimum detection concentration) described above, one can determine the number of molecules needed for analysis.

[0198] VI. Progress over time

[0199] As the tumor progresses, the amount of tumor DNA fragments in the plasma will increase as the tumor will release more DNA fragments (e.g., due to tumor growth, more necrosis, or higher vascularity). The increased amount of DNA fragments from the tumor tissue entering the plasma will increase the degree of imbalance in the plasma (e.g., in RHDO, the difference in the tag counts between the two haplotypes will increase). Moreover, as the amount of tumor DNA fragments increases, the number of regions in which aberrations exist can be more easily detected. For example, the amount of tumor DNA in a region can be too small for a chromosomal aberration to be detected. The reason for this is that when the tumor is small and releases only a small amount of cancer DNA fragments, there are not enough fragments to analyze, so a statistically significant difference cannot be established, making the aberration undetectable. Even when the tumor is small, more fragments can be available for analysis, but this can require a large amount of sample (e.g., a large amount of plasma).

[0200] Tracking the progression of a cancer can use the amount of aberrations in one or more regions (e.g., by imbalance or depth required), or the amount (number and / or size) of chromosomal regions exhibiting aberrations. In one example, if the amount of aberrations in one region (or multiple regions) is increasing faster than the amount of aberrations in other regions, then that region can be used as a preferred molecular marker to detect the cancer. This increase can be a result of the tumor growing larger and / or the region being amplified multiple times, releasing more DNA fragments. Moreover, one can also monitor changes in the post-surgery aberration values (e.g., the amount of aberrations or the number of regions exhibiting aberrations, or a combination thereof) to confirm whether the tumor was completely removed.

[0201] In various embodiments of the technology, determining the relative percent concentration of tumor DNA can be used to stage, prognose, or monitor the progression of a cancer. The progression determined can provide information about the current stage of the cancer and how quickly the cancer is growing or spreading. The "stage" of a cancer relates to all or a portion of the size of the tumor, histological appearance, presence / absence of lymph node metastasis, and presence / absence of distant metastasis. The "prognosis" of a cancer relates to assessing the probability of disease progression and / or the probability of survival from the cancer. Moreover, it also relates to the assessment of time during which the patient can not have clinical evolution or survival time. The "monitoring" of a cancer relates to seeing if the cancer is progressing (e.g., increasing in size, metastasizing to lymph nodes, or spreading to distant organs, i.e., metastasis). Moreover, monitoring can also relate to checking if the tumor is under control of a therapy. For example, if the therapy is effective, one can see a decrease in the size of the tumor, regression of metastasis or lymph node metastasis, improvement in the overall health of the patient (e.g., weight gain).

[0202] A. Determination of the relative percent concentration of cancer DNA

[0203] One method for tracking an increase in the amount of aberration in one or more regions is to determine the relative percentage concentration of cancer DNA in that region. The tumor can then be tracked over a period of time using changes in the relative percentage concentration of cancer DNA. This tracking can be used for diagnosis; for example, the first measurement can provide a background level (which may correspond to the general aberration level in a person), while subsequent measurements, if showing changes, indicate tumor growth (and therefore cancer). Furthermore, changes in the relative percentage concentration of cancer DNA can be used to evaluate the prognostic effect of treatment. In other embodiments of the technique, an increase in the relative percentage concentration of tumor DNA in plasma indicates a poorer prognosis or an increased tumor burden in the patient.

[0204] The relative percentage concentration of cancer DNA can be determined in several ways. For example, in tag counting, the difference between one haplotype and another (or one region compared to another). Another approach is to look at the depth prior to statistically significant differences (i.e., the number of fragments analyzed). An example of the former is that haplotype dose differences can be used to determine the relative percentage concentration of tumor-derived DNA in biological samples (e.g., plasma) by analyzing chromosomal regions with loss of heterozygosity.

[0205] Previous studies have shown a positive correlation between the amount of tumor-derived DNA and tumor burden in cancer patients (Lo et al. Cancer Res. 1999; 59:5452-5. and Chan et al. Clin Chem. 2005; 51:2192-5). Therefore, continuous monitoring of the relative percentage concentration of tumor-derived DNA in biological samples (e.g., plasma samples) via RHDO analysis can be used to monitor the progression of a patient's disease. For example, monitoring the relative percentage concentration of tumor-derived DNA in samples collected sequentially after treatment (e.g., plasma) can be used to determine the success of treatment.

[0206] Figure 16 The principle of measuring the relative percentage concentration of tumor-derived DNA in plasma by RHDO analysis according to an embodiment of the present invention is illustrated. The imbalance between two haplotypes is determined, and the degree of imbalance can be used to determine the relative percentage concentration of tumor DNA in the sample.

[0207] Hap I and Hap II represent two haplotypes in non-tumor tissues. Hap II is partially deleted in subregion 1610 in tumor tissues. Therefore, the Hap II-related fragments detected in plasma corresponding to the deleted region 1610 are contributed by non-tumor tissues. On the other hand, region 1610 in Hap I is present in both tumor and non-tumor tissues. Therefore, the difference between Hap I and Hap II readings represents the amount of tumor-derived DNA in plasma.

[0208] The following formula can be used to calculate the relative percentage concentration of tumor-derived DNA from the number of sequencing reads (tags) obtained from the deleted and non-deleted chromosomes for a chromosomal region affected by LOH: F = (N HapI -N HapII ) / N HapI x 100%, where N HapI is the number of sequencing reads corresponding to the allele on Hap I for heterozygous SNPs located in the chromosomal region affected by LOH; and N HapII is the number of sequencing reads corresponding to the allele on Hap II for heterozygous SNPs located in the chromosomal region affected by LOH.

[0209] The above formula is equivalent to defining p as the cumulative tag count for heterozygous loci located on the chromosomal region not including the deletion (Hap I), and q as the cumulative tag count for heterozygous loci located on the chromosomal region including the deletion (Hap II), and the relative percentage concentration of tumor DNA in the sample (F) is calculated as F = 1 - q / p. As shown in the example, the relative percentage concentration of tumor DNA is 14% (1 - 104 / 121). Figure 11

[0210] A plasma sample A with a certain percentage concentration of tumor DNA is collected from a HCC patient before and after tumor resection. Before tumor resection, N HapI is 30443 for the first haplotype of a given chromosomal region, and N HapII is 16221 for the second haplotype of the chromosomal region, which gives F as 46.7%. After tumor resection, N HapI is 31534, and N HapII is 31098, which gives F as 1.4%. This monitoring shows that the tumor resection is successful.

[0211] The degree of change in the size distribution of the circulating DNA can also be used to determine the relative percentage concentration of tumor DNA. In one embodiment, the exact size distribution of the tumor and non-tumor tissue DNA in the plasma can be determined, and then the measured size distribution that falls between the two known distributions can provide the percentage concentration of tumor DNA (e.g., using a linear model between the two statistical values of the size distribution of the tumor and non-tumor tissue). In addition, the size change can also be used for continuous monitoring. In one aspect, the change in size distribution is determined to be proportional to the percentage concentration of tumor DNA in the plasma.

[0212] ​In addition, the differences between the various regions can also be used in a similar manner, i.e., the non-specific haplotype detection method described above. In the tag count method, multiple parameters are used to monitor the progression of the disease. For example, the magnitude of the z-score for a region exhibiting chromosomal aberrations can be used to reflect the percentage concentration of tumor-derived DNA in a biological sample, such as plasma. The degree of over- or under-representation of a particular region is proportional to the percentage concentration of tumor-derived DNA in the sample or the degree or amount of copy number change in the tumor tissue. The magnitude of the z-score is a parameter that measures the degree of over- or under-representation of a particular chromosomal region in the sample as compared to a control subject. Thus, the magnitude of the z-score can reflect the relative percentage concentration of tumor DNA in the sample and, in turn, the tumor burden of the patient.

[0213] B. Number of Tracked Regions

[0214] As mentioned above, the number of regions exhibiting chromosomal aberrations can be used to screen for cancer and can also be used to monitor and prognosticate. For example, the monitoring can be used to determine the status of the current stage of cancer if the cancer recurs or if the treatment is working. As a tumor progresses, the genomic makeup of the tumor will be more degraded. To identify this continuous degradation, a method that tracks the number of regions (e.g., a predetermined region of 1 Mb) can be used to identify the progression of the tumor. In more advanced cancers, the tumor will have more regions exhibiting aberrations.

[0215] C. Methods

[0216] Figure 17 To illustrate a method for determining the progression of chromosomal aberrations in an organism using a biological sample comprising nucleic acid molecules according to an embodiment of the present application, a flowchart is shown. In one embodiment, at least some of the nucleic acid molecules are free. For example, the chromosomal aberrations can be from a malignant tumor or a pre-malignant lesion. In addition, the increase in aberrations can be due to the organism having more and more cells with chromosomal aberrations over time or due to the organism having a certain proportion of cells with an increasing amount of aberrations per cell. As a decrease, a treatment (e.g., surgery or chemotherapy) can remove or reduce the cells associated with the cancer.

[0217] In step 1710, one or more non-overlapping chromosomal regions of the organism are identified. Each chromosomal region comprises a plurality of loci. The regions can be identified by any suitable method, such as those described herein.

[0218] Steps 1720-1750 are performed for each time at multiple time points. Each time point corresponds to a different time at which a sample is obtained from the organism. The current sample is the sample that is analyzed during a given time. For example, a sample can be taken once a month for 6 months, and can be analyzed as soon as possible after the sample is obtained. Alternatively, the analysis can be performed after multiple measurements are taken over multiple time periods.

[0219] In step 1720, the current biological sample of the organism is analyzed to identify the location of the nucleic acid molecules in the reference genome of the organism. The location can be determined in any manner mentioned herein, for example by sequencing the fragments to obtain sequencing tags, and aligning the sequencing tags to the reference genome. In addition, the specific haplotype of the molecule can also be determined for haplotype-specific based methods.

[0220] Steps 1730-1750 are performed for each of the one or more chromosomal regions. When multiple regions are used, the implementation from section V can be used. In step 1730, based on the identified locations, the chromosomal region to which each set of nucleic acid molecules corresponds is identified. Each set of nucleic acids includes at least one nucleic acid molecule at each of a plurality of loci located in the chromosomal region. In one implementation, the set can be fragments that align to a specific haplotype of the chromosomal region, for example as described in the RHDO method above. In another implementation, the set can be any fragments that align to the chromosomal region, as described in the methods of section IV.

[0221] In step 1740, the computer system calculates a value for each set of nucleic acid molecules. The value defines a property of the set of nucleic acid molecules. The value can be any value mentioned herein. For example, the value can be the number of fragments in the set, or a statistical value of the size distribution of the fragments in the set. In addition, the value can be a normalized value, for example the number of tag counts for the region divided by the total number of tag counts for the sample, or divided by the number of tag counts for a reference region. In addition, the value can be a difference or ratio with another value (for example in the RHDO), thereby providing a differential property for the region.

[0222] In step 1750, the values are compared to a reference value to determine a classification of whether the first chromosomal region exhibits a deletion or amplification. The reference value can be any threshold or reference value described herein. For example, the reference value can be a threshold value determined for normal samples. For RHDO, the values can be the difference or ratio of the tag counts for the two haplotypes, and the reference value can be a threshold value for determining the presence of a statistically significant difference. As another example, the reference value can be the tag count or size of the other haplotype or region, and the comparison can include, but is not limited to, the difference or ratio (or a function of such values), and then determining whether the difference or ratio is greater than a threshold value. The reference value can be determined according to any suitable method and criteria, such as described herein.

[0223] In step 1760, the classifications of the various chromosomal regions at multiple time points are used to determine the nature of the chromosomal aberration in the organism. The process can be used to determine whether the organism has cancer, the stage of the cancer, and the prognosis of the cancer. Each of these determinations can involve a classification of the cancer, as described herein.

[0224] The classification of the cancer can be performed in a variety of ways. For example, the amount of aberrant regions can be counted and compared to a threshold value. In the case of regions, the classification can be a numerical value (e.g., the concentration of the tumor, and the reference value can be a value for a different haplotype or a different region), and the change in concentration can be determined. The change in concentration can be compared to a threshold value, and if a significant increase is determined to have occurred, a signal that a tumor is present is indicated.

[0225] VII. EXAMPLES

[0226] A. RHDO using SPRT

[0227] In this section, we show an example of the analysis of relative haplotype dosage (RHDO) using SPRT for a patient with hepatocellular carcinoma (HCC). In the tumor tissue of this patient, a deletion of one of the two copies of chromosome 4 was observed. This results in the loss of heterozygosity for SNPs on chromosome 4. The genome DNA of the patient, his wife, and his son were analyzed, and the genotypes of the three individuals were determined. The structural haplotypes of the patient were inferred from their genotypes. Massively parallel sequencing was performed, and sequencing reads with SNP alleles corresponding to the two haplotypes of chromosome 4 were identified and counted.

[0228] The equations and principles of RHDO and SPRT have been described above. In one embodiment, RHDO analysis can be programmed by a computer to detect a 10% difference in haplotype dosage, for example, in a DNA sample, where a 10% difference in dosage of one haplotype to another corresponds to a 10% concentration of tumor DNA. In other embodiments, the sensitivity of RHDO analysis can be set to detect 2%, 5%, 15%, 20%, 25%, 30%, 40%, and 50% of tumor-derived DNA in a DNA sample. The sensitivity of RHDO analysis can be adjusted using parameters that calculate the upper and lower threshold of the SPRT classification region. The adjustable parameters can be the desired level of detection (e.g., what percentage of tumor concentration should be detectable, which will affect the number of molecules analyzed) and the threshold for classification, for example, using the odds ratio (the ratio of the tag count of one haplotype to the tag count of the other haplotype).

[0229] In this RHDO analysis, the null hypothesis is that the two haplotypes of chromosome 4 are present at the same dosage. The alternative hypothesis is that the two haplotypes are present at more than a 10% difference in dosage in the biological sample (e.g., plasma). The number of sequencing reads with SNP alleles (corresponding to the two haplotypes) is compared statistically in the form of accumulated data from different SNPs for both hypotheses. When the accumulated data is sufficient to determine whether the two haplotypes are present at equal dosage or at least 10% different in a statistical manner, SPRT classification is performed. A typical SPRT classification region for the q arm of chromosome 4 is shown in Figure 18A The threshold of 10% is used here for illustrative purposes only. Other degrees of difference (e.g., 0.1%, 1%, 2%, 5%, 15%, or 20%) can also be detected. Generally, the lower the degree of difference one wants to detect, the more DNA molecules one needs to analyze. Conversely, the higher the degree of difference one wants to detect, the fewer DNA molecules one needs to analyze to achieve a statistically significant difference. For this analysis, the odds ratio is used for SPRT, but other parameters such as z-score or p-value can be used.

[0230] In the plasma samples taken at diagnosis of HCC patients, there were 76 and 148 successful RHDO classifications for the p and q arms of chromosome 4, respectively. All of the RHDO classifications showed an imbalance in haplotype dosage in the plasma samples taken at diagnosis. As a comparison, the plasma samples of patients taken after surgical resection of the tumor were also analyzed, as shown in Figure 18BThe results are shown in Figure 6. For the post-treatment sample, there were 4 and 9 successful RHDO classifications for the p and q arms of chromosome 4, respectively. All 4 RHDO classifications showed that there was no observable haplotype dosage imbalance corresponding to tumor DNA concentrations greater than 10% in the plasma sample. Of the 9 RHDO classifications for chromosome 4q, 7 showed a lack of haplotype dosage imbalance, while 2 showed an imbalance. The number of RHDO blocks showing a dosage imbalance corresponding to tumor DNA concentrations greater than 10% was significantly reduced after tumor resection, indicating that the size of the chromosomal region showing a dosage imbalance was significantly smaller in the post-treatment sample than in the pre-treatment sample. These results indicate that the relative percentage concentration of tumor DNA in the plasma is reduced after surgical resection of the tumor.

[0231] When compared to non-haplotype specific methods, RHDO analysis provides a more accurate estimate of the relative percentage concentration of tumor DNA and is particularly suitable for monitoring the progression of aberrations. Thus, one can expect that a case of disease progression would be manifested by an increase in the percentage concentration of tumor DNA in the plasma; whereas, if a patient's disease is stable or their tumor is regressing or shrinking in size, the percentage concentration of tumor DNA in the plasma would decrease.

[0232] B. Targeted Analysis

[0233] In some alternative embodiments, target sequence capture techniques can be combined with universal sequencing of DNA fragments. This approach is also referred to herein as targeted analysis. One embodiment of this approach is to use a liquid phase capture system (such as the Agilent SureSelect system, the Illumina TruSeq Custom Enrichment Kit (illumina.com / applications / sequencing / targeted_resequencing.ilmn), or by the MyGenostics GenCap Custom Enrichment system (mygenostics.com / )) or a microarray-based capture system (such as the Roche NimbleGene system) to preferentially select fragments. Certain regions are preferentially captured, although some other regions can be captured as well. Such an approach can allow such regions to be analyzed at a higher depth (e.g., more fragments can be sequenced or analyzed using digital PCR) and / or at a lower cost. Greater depth can increase sensitivity in the region. Other enrichment methods can be implemented based on the size and methylation pattern of the fragments.

[0234] Therefore, an alternative approach to the analysis of DNA samples on a whole genome basis lies in targeted analysis of regions of interest in order to detect common chromosomal aberrations. As the analysis method focuses primarily on regions where chromosomal aberrations are potentially present, or regions that have specific characteristic changes that correspond to specific tumours, or regions that have specific clinical importance, targeted analysis methods can potentially improve the cost of the method. Examples of changes such as those described above include those that occur early in the formation of tumours of specific cancer types (for example in HCC, amplification of 1q and 8q and deletion of 8q are early chromosomal changes - van Malenstein et al. Eur J Cancer 201 1 ; 47: 1789-97); or changes that are associated with good or poor prognosis (for example increased gain at 6q and 17q and loss at 6p and 9p are observed during tumour progression, and in colorectal cancer patients, the presence of LOH at 18q, 8p and 17p is associated with poorer survival - Westra et al. Clin Colorectal Cancer 2004; 4: 252-9); or which can predict response to therapy (for example the presence of increased 7p is predictive of response to tyrosine kinase inhibitors in patients with epidermal growth factor receptor mutations - Yuan et al. J Clin Oncol 201 1 ; 29: 3435-42). Other examples of genomic regions that are altered in cancer can be found in a number of online databases (for example the Cancer Genome Anatomy Project database (cgap.nci.nih.gov / Chromosomes / RecurrentAberrations) and Atlas of Genetics and Cytogenetics in Oncology and Haematology (atlasgeneticsoncology.org / / Tumors / Tumorliste.html)). In contrast, in a non-targeted whole genome approach, regions where chromosomal aberrations are unlikely to occur are analysed to the same extent (depth) as regions where aberrations are potentially present.

[0235] We used a targeted analysis strategy to analyze plasma samples obtained from 3 HCC patients and 4 healthy control subjects. The SureSelect capture system from Agilent was used to achieve the enrichment of targets (Gnirke et al. Nat. Biotechnol 2009. 27: 182-9). The SureSelect system was chosen as an example of a feasible target enrichment technology. Other liquid phase (Illumina TruSeq Custom Enrichment system) or solid phase (e.g. Roche-Nimblegen system) target capture systems as well as amplification-based target enrichment systems (e.g. QuantaLife system and RainDance system) can also be used. The capture probes were designed to be located on regions of chromosomes that are commonly and rarely shown to be aberrant in HCC. After target capture, the samples were then sequenced in a flowcell of an Illumina GAIIx analyzer in a format of one lane per DNA sample. Regions with very low probability of amplification and deletion were used as reference to compare with regions where amplification and deletion are relatively frequently present.

[0236] In Figure 19 , common chromosomal aberrations found in HCC were shown (the figure was adapted from Wong et al (Am J Pathol 1999; 154: 37-43)). The lines on the right side of the chromosome pattern figure indicate chromosomal gains, while the lines on the left side indicate chromosomal losses of each patient sample. The thick lines indicate high levels of gains. The rectangles indicate the locations of the target capture probes.

[0237] Targeted tag count analysis

[0238] For the detection of chromosomal aberrations, we first calculated the normalized tag counts for the potential aberrant region and the reference region. Next, the normalized tag counts were corrected for the GC content of the region as previously described by Chen et al (PLoS One 2011; 6: e21791). In the current example, the p-arm of chromosome 8 was chosen as the potential aberrant region and the q-arm of chromosome 9 as the reference region. The tumor tissues of 3 HCC patients were analyzed for chromosomal aberrations using the Affymetrix SNP 6.0 microarray. The changes in the chromosomal dosage of 8p and 9q in the tumor tissues of the 3 patients are shown below. Patient HCC013 had a loss of 8p with no change in 9q. Patient HCC027 had a gain of 8p with no change in 9q. Patient HCC023 had a loss of 8p with no change in 9q.

[0239] Next, the ratio of the normalized tag counts between chr 8p and 9q was calculated for 3 HCC patients and 4 healthy control subjects using targeted analysis. Figure 20A Results of the normalized tag count ratios for HCC and healthy patients are shown. For the cases of HCC013 and HCC023, a decrease in the normalized tag count ratio between 8p and 9q was observed. This is consistent with the loss of chromosome 8p found in the tumor tissue. For the case of HCC027, an increase in the ratio was observed, and this increased ratio is consistent with the gain of chromosome 8p in the tumor tissue of this case. The dashed lines indicate a shift of plus or minus two standard deviations from the mean of the 4 normal control subjects, respectively.

[0240] Targeted size analysis

[0241] In the previous section, we described the principle of identifying cancer-related alterations by determining the size distribution of plasma DNA fragments in cancer patients. In addition, a target enrichment method can be used to detect the size alterations. For the 3 HCC cases (HCC 013, HCC027 and HCC023), the size of each sequenced DNA fragment was determined after aligning the sequencing reads to the human reference genome. The size of the sequenced DNA fragment was deduced by the coordinates of the two outermost nucleotides. In other embodiments, the entire length of the DNA fragment is sequenced, and the size of the fragment can then be determined directly from the sequencing length. The size distribution of the DNA fragments aligned to chromosome 8p was compared to the size distribution of the DNA fragments aligned to chromosome 9q. For detecting the difference in the size distribution of the DNA of the two populations, first the proportion of DNA fragments shorter than 150 bp was determined for each population in the current example. In other embodiments, other cutoff values can be used, such as 80 bp, 110 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 160 bp and 170 bp. ΔF = F 8p - Q 9q where F 8p is the proportion of DNA fragments aligned to chromosome 8p and shorter than 150 bp; and F 9q is the proportion of DNA fragments aligned to chromosome 9q and shorter than 150 bp.

[0242] Since a shorter DNA fragment size distribution results in a higher proportion of fragments shorter than the cutoff value (in the current example, shorter than 150 bp), a higher (more positive) value of ΔF indicates that the DNA fragments aligned to chromosome 8p have a shorter distribution relative to those aligned to chromosome 9q. In contrast, a lower (or more negative absolute value) result indicates that the DNA fragments aligned to chromosome 8p have a longer distribution relative to those aligned to chromosome 9q.

[0243] Figure 20B Results of size analysis after target enrichment and massively parallel sequencing are shown for 3 HCC patients and 4 healthy control subjects. In the 4 healthy control subjects, the positive ΔQ indicates that DNA fragments mapping to chromosome 8p have a slightly shorter size distribution than those mapping to chromosome 9q. The dashed lines indicate the interval corresponding to ΔQ values obtained within 2 standard deviations from the mean for the 4 control subjects. The ΔQ values for HCC cases 013 and HCC 023 are more than 2 standard deviations below the mean for the control subjects. Both of these cases have a deletion of chromosome 8p in the tumor tissue. Deletion of 8p in the tumor can result in a reduced contribution of tumor-derived DNA to the plasma with respect to this chromosomal region. Since tumor-derived DNA is shorter than DNA derived from non-tumor tissue in the peripheral blood, this can result in a significantly longer size distribution of plasma DNA fragments mapping to chromosome 8p. This is consistent with the lower (more negative) ΔQ values in these two cases. In contrast, amplification of 8p in the case of HCC 027 can result in a significantly shorter distribution of DNA fragments mapping to this region. Therefore, a higher proportion of plasma DNA fragments mapping to 8p are considered to be relatively short. This is consistent with the observed case of HCC 027 having a positive ΔQ value with an absolute value greater than the healthy control subjects.

[0244] C. Multiple regions for detecting tumor-derived chromosomal aberrations

[0245] Chromosomal aberrations, including deletions and amplifications of certain chromosomal regions, are commonly detected in tumor tissues. Characteristic patterns of chromosomal aberrations are observed in different types of cancer. Here, we use multiple examples to illustrate different approaches for detecting these cancer-associated chromosomal aberrations in the plasma of cancer patients. In addition, our approach is also suitable for cancer screening, monitoring of disease progression, and response to therapy. Samples obtained from one HCC patient and two nasopharyngeal carcinoma (NPC) patients were analyzed. For the HCC patient, venous blood samples were collected before and after surgical removal of the tumor. For the two NPC patients, venous blood samples were collected at the time of diagnosis. In addition, plasma samples from one chronic hepatitis B carrier and one subject with detectable Epstein-Barr virus DNA in plasma were analyzed. Both of these subjects did not have any cancer.

[0246] Microarray analysis was used to perform detection of tumor-derived chromosomal aberrations. Specifically, DNA extracted from blood cells and tumor samples of HCC patients were analyzed using the Affymetrix SNP 6.0 microarray system. Genotypes of blood cells and tumor tissues were determined using Affymetrix Genotyping Console v4.0. Chromosomal aberrations, including gains and deletions, were determined using the Birdseed v2 algorithm based on signal intensities of different alleles of SNPs on the microarray and copy number alteration (CNV) probes.

[0247] Count-based analysis

[0248] To perform sequencing tag count analysis in plasma, 10 milliliters of venous blood was collected from each subject. For each blood sample, plasma was collected after the sample was centrifuged. DNA was extracted from 4-6 mL of plasma using the QIAmp blood mini Kit (Qiagen). Plasma DNA libraries were constructed as previously described (Lo YMD. Sci Transl Med 2010, 2: 61ra91) and then subjected to massively parallel sequencing using the Illumina Genome Analyzer platform. Paired-end sequencing was performed for the analysis of plasma DNA. Each molecule was sequenced at each end of the two ends (50 bp), so each molecule totaled 100 bp. The two ends of each sequence were aligned to the human genome (Hgl8 NCBI.36 downloaded from UCSC genome.ucsc.edu) using the SOAP2 program (soap.genomics.org.cn / ) (Li R et al. Bioinformatics 2009, 25: 1966-7).

[0249] The genome is then divided into a number of 1 megabase (1 Mb) chromosomal segments, and the number of sequencing reads aligned to each 1 Mb chromosomal segment is determined. Next, the tag counts for each segment are corrected using a locally weighted regression scatterplot smoothing (LOESS)-based algorithm according to the GC content of each chromosomal segment (Chen E et al. PLoS One 2011, 6: e21791). The purpose of this correction is to minimize sequencing-related quantitative bias due to differences in GC content among different genomic segments. The division into 1 Mb segments is for illustrative purposes. Other segment sizes can also be used, e.g., 2 Mb, 10 Mb, 25 Mb, 50 Mb, etc. In addition, the segment size can also be chosen based on the genomic characteristics of the particular tumor of the particular patient, as well as the particular type of tumor in general. Furthermore, if the sequencing method can show lower GC bias, e.g., the Helicos system (www.helicosbio.com) or the Pacific Biosciences Single Molecular Real-Time system (www.pacificbiosciences.com), the GC correction step can be omitted, e.g., for a single molecular sequencing technology.

[0250] In a previous study, we sequenced 57 plasma samples obtained from subjects who did not have cancer. The results of these plasma sequencing were used to determine the reference range of tag counts for each 1 Mb segment. For each 1 Mb segment, the mean and standard deviation of the tag counts of the 57 individuals were determined. Next, the results of the study subjects were expressed as z-scores, which were calculated using the following equation: z-score = (amount of sequencing tags for the case - mean) / S.D., where the mean value is the mean value of sequencing tags aligned to a particular 1 Mb segment for the reference samples; S.D. is the standard deviation of sequencing tags aligned to a particular 1 Mb segment for the reference samples.

[0251] Figures 21-24 The results of sequencing tag count analysis for 4 study subjects are shown. The 1 Mb segments are shown on the edges of the figure. The chromosome numbers and the chromosome ideogram (outermost circle) are adjusted in a clockwise direction from pter to qter (centromere is shown in yellow). The results for each subject are shown in a different color. The results are shown as z-scores, which are calculated using the following equation: z-score = (amount of sequencing tags for the case - mean) / S.D., where the mean value is the mean value of sequencing tags aligned to a particular 1 Mb segment for the reference samples; S.D. is the standard deviation of sequencing tags aligned to a particular 1 Mb segment for the reference samples. Figure 21In the center, the inner circle 2101 shows the regions of aberration (deletion or amplification) determined by analysis of the tumor. The inner circle 2101 is shown in 5 levels. The levels go from -2 (the innermost line) to +2 (the outermost line). A value of -2 indicates loss of both copies of a chromosome for the corresponding region. A value of -1 indicates loss of one of the two copies of a chromosome for a certain chromosomal region. A value of 0 indicates no gain or loss of a chromosome. A value of +1 indicates gain of one copy of a chromosome, and +2 indicates gain of two copies of a chromosome.

[0252] The middle circle 2102 shows the results of the plasma analysis. As one can see, the middle circle 2102 corresponds to the aberration results in the pre-treatment plasma and matches the aberration pattern of the inner circle. The middle circle 2102 is for more levels of lines, but the progression is the same. The outer circle 2103 shows the data points obtained from analyzing the plasma after treatment, and these data points are in gray (proving under-presentation / presentation - no aberration).

[0253] Chromosomal regions that are over-presented in the plasma (z-score > 3) are represented by green dots 2110. Regions that are under-presented in the plasma (z-score < -3) are represented by red dots 2120. Regions that have no significant chromosomal aberration detected in the plasma (z-score between -3 and 3) are represented by gray dots. The over-presentation / under-presentation is normalized to the total number of counts. In case of PCR amplification prior to sequencing, the normalization can take into account correction for GC bias.

[0254] Figure 21 A Circos plot of an HCC patient according to an embodiment of the application is shown, depicting the data obtained from the counts of sequencing tags of plasma DNA. The circles from the inside out represent respectively: chromosomal aberrations detected in the tumor tissue by microarray analysis (red and green for deletions and amplifications, respectively); z-score analysis of a plasma sample obtained prior to surgical resection of the tumor; and z-score analysis of a plasma sample obtained 1 month after resection. Prior to tumor resection, the chromosomal aberrations detected in the plasma match exactly those identified in the tumor tissue by microarray analysis. After tumor resection, most of the cancer-related chromosomal aberrations disappear in the plasma. These data reflect the value of such a method for detecting disease progression and treatment efficacy.

[0255] Figure 22 A count analysis of sequencing tags performed on a plasma sample of a chronic HBV carrier not suffering from HCC according to an embodiment of the application is shown. In contrast to the HCC patient ( Figure 21 ), cancer-related chromosomal aberrations are not detected in the plasma of this HBV carrier. These data reflect the value of this method for cancer screening, diagnosis and monitoring.

[0256] Figure 23 This paper illustrates a counting analysis of sequencing tags in plasma samples from cancer patients with stage III NPC, according to an embodiment of the invention. Chromosomal aberrations were detected in the plasma samples obtained prior to treatment. Specifically, significant aberrations were identified in chromosomes 1, 3, 7, 9, and 14.

[0257] Figure 24 This illustration shows a sequencing tag counting analysis of plasma samples from cancer patients with stage IV NPC according to an embodiment of the present invention. Chromosomal aberrations were detected in the plasma samples obtained before treatment. When compared with patients with stage III cancer (… Figure 23 Compared to patients with [specific disease type], more chromosomal aberrations were detected. Furthermore, the sequencing tag counts deviated more from the control mean, i.e., the z-scores deviated more from 0 (positive or negative). The increased number of chromosomal aberrations and the greater degree of deviation in sequencing tag counts compared to controls reflect a more severe degree of genomic alteration in advanced disease, thus highlighting the value of such methods for cancer staging, prognosis, and surveillance.

[0258] Size-based analysis

[0259] Previous studies have shown that the size distribution of DNA derived from tumor tissue is shorter than that of DNA derived from non-tumor tissue (Diehl F et al. Proc Natl Acad Sci USA 2005, 102(45):16368-73). In previous studies, we summarized a method for detecting plasma haplotype imbalances through plasma DNA size analysis. Here, we use sequencing data from HCC patients to further illustrate this method.

[0260] For illustrative purposes, we identified two regions for size analysis. In one region (chromosome 1 (chr1); coordinates: 159,935,347 to 167,219,158), a duplicate of one of the two homologous chromosomes was detected in tumor tissue. In another region (chromosome 10 (chr10); coordinates: 100,137,050 to 101,907,356), a deletion (LOH) of one of the two homologous chromosomes was detected in tumor tissue. In addition to determining which haplotype the sequencing fragment originated from, the size of the sequencing fragment was determined bioinformatically using the coordinates of the outermost nucleotide of the sequencing fragment in the reference genome. Next, the size distribution of the fragments corresponding to each of the two haplotypes was determined.

[0261] For the LOH region on Chr 10, one haplotype (i.e., the deleted haplotype) was detected in the tumor tissue. Therefore, all plasma DNA fragments that map to this deleted haplotype are derived from non-cancerous tissue. On the other hand, fragments that map to the haplotype that is not deleted in the tumor tissue (non-deleted haplotype) can be derived from either the tumor or non-tumor tissue. Since the size distribution of tumor-derived DNA is shorter, we predict that the size distribution of fragments corresponding to the non-deleted haplotype is shorter than the size distribution of fragments corresponding to the deleted haplotype. The difference between the two size distributions can be determined by plotting the cumulative frequency of fragments against the size of the DNA fragments. The DNA population with the shorter size distribution will have a relatively larger number of short DNA fragments, and therefore, the cumulative frequency corresponding to the DNA population with the shorter size distribution will increase more rapidly in the interval corresponding to the shorter fragments in the size range.

[0262] Figure 25 The relationship between the cumulative frequency of plasma DNA and size for the region exhibiting LOH in the tumor tissue is shown according to an embodiment of the application. The X-axis is the size of the fragments in base pairs. The Y-axis is the percentage of fragments with size less than the value on the X-axis. Sequences derived from the non-deleted haplotype have a more rapid increase and a higher cumulative frequency for sizes less than 170 bp compared to sequences derived from the deleted haplotype. This indicates that short DNA fragments derived from the non-deleted haplotype are more abundant. This is consistent with the above prediction due to the size distribution of tumor-derived short DNA from the non-deleted haplotype.

[0263] In one embodiment, the difference in size distribution can be measured by the difference in the cumulative frequency of the two DNA populations. We define AQ as the difference in the cumulative frequency of the two populations. AQ = Q 非删除的 - Q 删除的 , Q 非删除的 represents the cumulative frequency of sequenced DNA fragments derived from the non-deleted corresponding haplotype; and Q 删除的 represents the cumulative frequency of sequenced DNA fragments derived from the deleted corresponding haplotype.

[0264] Figure 26 The relationship between AQ and the size of sequenced plasma DNA for the LOH region is shown. According to an embodiment of the application, AQ reaches 0.2 at a size of 130 bp. This indicates that using 130 bp as a cutoff for defining short DNA is optimal for use in the equation described above. Using this cutoff, the amount of short DNA is 20% higher in the population derived from the non-deleted haplotype when compared to the population derived from the deleted haplotype. This percentage difference (or similar derived value) is then compared to a threshold value derived from individuals who do not have cancer.

[0265] With respect to regions with chromosomal amplification, in tumor tissues, one haplotype is duplicated (amplified haplotype). Since additional amounts of tumor-derived short DNA molecules from this amplified haplotype are released into the plasma, the size distribution of fragments from the amplified haplotype is shorter than that from the non-amplified haplotype. Similar to the LOH approach, the difference in size distribution can be measured by plotting the cumulative frequency of DNA fragments against their size. The population of DNA with shorter size distribution has more short DNA, and thus its cumulative frequency increases more rapidly in the interval corresponding to short fragments.

[0266] Figure 27 The cumulative frequency of plasma DNA against the size of DNA fragments corresponding to regions with chromosomal duplication in tumor tissues is shown according to embodiments of the application. Sequences from the amplified haplotype have a more rapid increase and a higher cumulative frequency corresponding to sizes less than 170 bp than sequences from the non-amplified haplotype. This indicates that there are more short DNA fragments from the amplified haplotype. Since a large amount of tumor-derived short DNA is derived from the amplified haplotype, the above is consistent with the prediction shown below.

[0267] Similar to the LOH approach, the difference in size distribution can be quantified by the difference in the cumulative frequency of DNA molecules from the two populations. We define AQas the difference in the cumulative frequency of the two populations. AQ= Q 扩增的 - Q 非扩增的 , Q 扩增的 represents the cumulative frequency of sequenced DNA fragments from the amplified haplotype; and Q 非扩增的 represents the cumulative frequency of sequenced DNA fragments from the non-amplified haplotype.

[0268] Figure 28 The relationship between AQand the size of sequenced plasma DNA for amplified regions is shown according to embodiments of the application. According to embodiments of the application, AQreaches 0.08 at a size of 126 bp. This indicates that using 126 bp as a cutoff for defining short DNA, there are 8% more short DNA molecules in the population from the amplified haplotype than in the population from the non-amplified haplotype.

[0269] D. Other techniques

[0270] In other embodiments, sequence-specific techniques can be used. For example, oligonucleotides can be designed to hybridize to fragments of a particular region. The oligonucleotides can then be counted in a manner similar to sequencing tag counts. This approach can be used to detect cancers that exhibit particular aberrations.

[0271] VIII. Computer System

[0272] Any of the computer systems mentioned herein can utilize any suitable number or combination of subsystems. Examples of such subsystems are shown in the computer apparatus 900 in Figure 9 some embodiments, a computer system includes a single computer apparatus, where the subsystems can be departments of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses (each being a subsystem) and internal components.

[0273] Figure 29 The illustrated subsystems are interconnected via a system bus 2975. Other subsystems such as a printer 2974, a keyboard 2978, a mouse 2977, and a monitor 2976 can also be connected to the computer system via any of a variety of connections, such as a parallel port, a serial port, or an external interface 2981 (e.g., Ethernet, Wi-Fi, etc.). The computer system 2900 can be connected to a network (such as the Internet) via an external interface 2981. The interconnection via system bus allows the central processor 2973 to communicate with each subsystem and for the subsystems to communicate with each other. The internal communication can be carried out through either a shared system bus or a different type of internal communication known to persons of ordinary skill in the art. The system memory 2972 and / or the hard disk 2979 can be considered computer readable media. Any of the values mentioned herein can be output by one component to another component and to a user.

[0274] A computer system can include a plurality of the same components or subsystems, e.g., connected together by an external interface 2981 or by an internal interface. In some embodiments, computer systems, subsystems, or apparatuses can communicate over a network. In such instances, one computer can be used as a client to another computer, which can be used as a server, where each computer has portions stored as computer readable media. The client and server can each include multiple systems, subsystems, or components.

[0275] It should be understood that any of the elements of the application can be implemented in whole or in part by hardware and / or software (including firmware, resident software, micro-code, etc.). Furthermore, unless specified otherwise, any of the components of the application can be used as either hardware or software, including resident software, micro-code, circuitry, and the like. Any of the elements of the application can be used alone or together in any combination of one or more of the elements of the application. It should further be understood that any of the components of the application can be used in any combination of one or more of the elements of the application. It should further be understood that any of the components of the application can be used in any combination of one or more of the elements of the application.

[0276] Any of the software components or functions described in this application can be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C++ or Perl. The software code can be stored as a series of instructions or commands or in any non-transitory computer-readable medium for use by or in connection with an instruction execution system such as, for example, a processor-based system. Non-transitory computer-readable medium can include, but is not limited to, any type of media for storing or transmitting software code including removable media and non-removable media, volatile media and non-volatile media, and transmission media. The computer-readable medium can be any combination of such storage or transmission media.

[0277] Further, such programs can also be encoded and transmitted using carrier signals adapted to carry the program code using any of a variety of communications protocols, including Internet. Thus, a computer readable medium comprising a computer program product can be created using a data signal encoded with such program code. The computer readable medium can be of any consortium of hardware or software products by or in connection to which the program code is executed, provided that the computer program product includes any of the embodiments of the application as described herein. The computer system can include a monitor, printer or other suitable display for providing any of the results mentioned herein to a user.

[0278] Any of the methods described herein can be totally or partially performed with a computer system including, for example, a processor, a storage medium, an input device and an output device. For example, various embodiments of the present application can be implemented by a computer system that includes a processor, a storage medium, an input device and an output device. The storage medium can include a memory device such as a hard drive, a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc, a flash drive, a solid state drive, or any other suitable memory device. The storage medium can store one or more computer programs, including one or more of the software components of the present application. The input device can include a keyboard, a mouse, a trackball, a microphone, a camera, a scanner, a touch screen, a touchpad, a pen, a game controller, a voice recognition device, a bar code reader, or any other suitable input device. The output device can include a monitor, a printer, a speaker, a holographic display, a projector, a tactile / binary output device, or any other suitable output device. Although the computer system is depicted as having a storage medium, a processor and an input and output device, in other embodiments, the computer system can not have any of the above components.

[0279] The specific details of particular embodiments can be combined in any suitable manner without departing from the spirit and scope of the present application. However, other embodiments of the present application can be directed to specific embodiments relating to each individual aspect or a specific combination of aspects.

[0280] In order to illustrate and describe the application, the foregoing description has presented exemplary embodiments of the application. No person skilled in the art will, on the basis of the foregoing description, make modifications and variations in the form and details of the application without departing from the spirit and scope of the application. The embodiments are selected and described in order to best explain the principles of the application and its practical application, thereby enabling others skilled in the art to best exploit the application in various embodiments and with various modifications as suitable for the particular use contemplated.

[0281] Unless specifically noted otherwise, the description "a" or "an," or "the" means "one or more." All patents, patent applications, publications, and descriptions mentioned above are incorporated herein by reference in their entirety for all purposes. None is admitted to be prior art.

Claims

1. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to implement a method of identifying one or more chromosomal aberrations of an organism from a biological sample of a subject, the method comprising: (a) identifying a chromosomal region of a reference genome of an organism, the chromosomal region comprising a plurality of loci; (b) for each of a plurality of nucleic acid molecules in a biological sample of the subject, identifying a location of the nucleic acid molecule in the reference genome; (c) identifying, based on the identified locations, a first set of nucleic acid molecules as being from the chromosomal region; (d) calculating a first value for the first set of nucleic acid molecules, the first value defining a property of the first set of nucleic acid molecules; (e) comparing the first value to a first threshold value, the first threshold value being offset from a first normal value by a first amount; (f) as a result of the first value exceeding the first threshold value, identifying a second set of nucleic acid molecules as being from a sub-region within the chromosomal region, wherein the second set of nucleic acid molecules is a subset of the first set of nucleic acid molecules; (g) calculating a second value for the second set of nucleic acid molecules, the respective value defining a property of the second set of nucleic acid molecules; (h) comparing the second value to a second threshold value, the second threshold value being offset from a second normal value by a second amount, the second amount being of a higher order than the first amount; (i) identifying, based on the comparison of the second value to the second threshold value, whether the sub-region of the chromosomal region exhibits aberration.

2. The computer readable medium of claim 1, wherein the method further comprises repeating steps (a)-(i) for each of a plurality of chromosomal regions within the reference genome.

3. The computer readable medium of claim 2, wherein the plurality of chromosomal regions cover the reference genome.

4. The computer readable medium of claim 2, wherein the plurality of chromosomal regions are non-overlapping.

5. The computer readable medium of claim 2, wherein each of the plurality of chromosomal regions is of a predetermined length.

6. The computer readable medium of claim 2, wherein each of the plurality of chromosomal regions is of substantially equal length.

7. The computer readable medium of claim 2, wherein: the biological sample comprises cell-free DNA derived from normal cells and from cancer-related cells, and each chromosomal region comprises a first plurality of loci that are heterozygous in the normal cells.

8. The computer readable medium of claim 7, wherein the cancer-related cells exhibit loss of heterozygosity.

9. The computer readable medium of claim 1, wherein the first threshold value and / or the second threshold value is derived from one or more healthy organisms or from one or more regions that do not exhibit loss or amplification.

10. The computer readable medium of claim 1, wherein the biological sample comprises cell-free DNA derived from normal cells and from cancer-related cells, wherein the method further comprises: determining a first haplotype and a second haplotype for the chromosomal region in the normal cells.

11. The computer readable medium of claim 10, wherein the method further comprises: identifying a first subset of nucleic acid molecules from the first set of nucleic acid molecules as being from the first haplotype; and identifying a second subset of nucleic acid molecules from the first set of nucleic acid molecules as being from the second haplotype. ​ 12. The computer readable medium of claim 11, wherein the first value is a difference or a ratio between a first subpopulation value of the first subpopulation of nucleic acid molecules and a second subpopulation value of the second subpopulation of nucleic acid molecules.

13. The computer readable medium of claim 12, wherein the method further comprises: calculating a ratio of the first subpopulation value and the second subpopulation value to determine a fractional concentration of cancer DNA in the biological sample; and using the fractional concentration to diagnose, stage, predict, or monitor progression of cancer in the subject.

14. The computer readable medium of claim 12, wherein the first subpopulation value is a first statistical value of a size distribution of nucleic acid molecules from the first group and the second subpopulation value is a second statistical value of a size distribution of nucleic acid molecules from the second group.

15. The computer readable medium of claim 12, wherein the first subpopulation value is a number of nucleic acid molecules of the first subpopulation and the second subpopulation value is a number of nucleic acid molecules of the second subpopulation.

16. The computer readable medium of claim 1, wherein the first value corresponds to a statistical value of a size distribution of nucleic acid molecules from the chromosomal region.

17. The computer readable medium of claim 16, wherein the statistical value is a proportion of nucleic acid molecules shorter than a cutoff size.

18. A system comprising: the computer readable medium of any one of claims 1-17; and one or more processors for executing instructions stored on the computer readable medium.

Citation Information

Patent Citations

  • Eccentrically-adjustable cam for scales

    US1636873A

  • Diagnosing fetal chromosomal aneuploidy using massively parallel genomic sequencing

    US20090029377A1

  • Fetal Genomic Analysis From A Maternal Biological Sample

    US20110105353A1

  • Size-based genomic analysis

    US20110276277A1

  • Methods for identifying DNA copy number changes

    US20040157243A1