Detection of genetic or molecular distortions associated with cancer

By analyzing chromosomal imbalances in cell-free DNA fragments, the problem of low accuracy and sensitivity in existing cancer screening technologies has been solved, enabling efficient cancer diagnosis and monitoring.

CN121768657APending Publication Date: 2026-03-31THE CHINESE UNIVERSITY OF HONG KONG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2011-11-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing cancer screening technologies have low accuracy and sensitivity, and often require high doses of radiation, making them ineffective for cancer screening, prognosis, and monitoring.

Method used

By analyzing cell-free DNA fragments, we can identify chromosomal imbalances in tumors caused by chromosomal fragment deletions and amplifications. By utilizing chromosomal regions from multiple loci for calculations, we can achieve cancer diagnosis, screening, and prognostic monitoring.

Benefits of technology

It improves the efficiency and accuracy of cancer diagnosis and monitoring, enabling the tracking of imbalances in chromosomal regions at different time points, and providing monitoring of cancer screening, prognosis, and treatment effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768657A_ABST
    Figure CN121768657A_ABST
Patent Text Reader

Abstract

The present invention provides systems, instruments, and methods for determining genetic or molecular distortions in a biological sample from an organism. Biological samples including free DNA fragments are analyzed to identify imbalances present in chromosomal regions, e.g., due to tumor chromosome deletion and / or amplification. Multiple loci are used for analysis of individual chromosomal regions. Such imbalances can then be used to diagnose (screen) cancer as well as prognosis of cancer patients, or to detect a patient's pre-exacerbation health status or to monitor changes in a patient's pre-exacerbation health status. Diagnosis, screening, prognosis, and monitoring can be provided using the severity of genomic imbalances, as well as the number of regions of imbalance. Systematic analysis of non-overlapping chromosome segments can provide a universal cancer screening means. Furthermore, the patient can be examined at different points in time to track one or more chromosomal regions and the severity of each of the numerous chromosomal regions and the number of chromosomal regions exhibiting chromosomal aberrations, therefore, cancer screening, prognosis diagnosis and lesion process monitoring (for example, after treatment) can be carried out.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 61 / 418,391, filed November 30, 2010, entitled “DETECTION OF GENETIC ABERRATIONS ASSOCIATED WITH CANCER,” and U.S. Provisional Application No. 61 / 529,877, filed August 31, 2011, entitled “DETECTION OF GENETIC OR MOLECULAR ABERRATIONS ASSOCIATED WITH CANCER,” and is a non-provisional application thereof, the entire contents of which are incorporated herein by reference for all purposes.

[0003] This application relates to two jointly owned U.S. patent applications filed November 5, 2010, entitled “Size-Based Genomic Analysis”, namely U.S. Patent Application No. 12 / 940,992 (U.S. Publication 2011 / 0276277) (Lo et al.) (Attorney’s File No. 80015-794101 / 006610US) and U.S. Patent Application No. 12 / 940,993 (U.S. Publication 2011 / 0105353) (Lo et al.) (Attorney’s File No. 80015-794103 / 006710US), entitled “Fetal Genomic Analysis From A Maternal Biological Sample”, both filed November 5, 2010, the disclosures of which are incorporated herein by reference in their entirety. Background Technology

[0004] Cancer is a common disease affecting many people. It is often not discovered until severe symptoms appear. While some cancer screening technologies exist to determine the likelihood of a patient having cancer for common types, their reliability and accuracy are often low, or they require patients to be exposed to high doses of radiation. For many other types of cancer, there are currently no effective screening technologies.

[0005] Deletion of heterozygosity (LOH) at specific loci has been detected in peripheral blood DNA from patients with lung cancer and head and neck cancer (Chen XQ, et al. Nat Med 1996; 2: 1033-5; Nawroz H, et al. Nat Med 1996; 2:1035-7). However, the relatively small amount of LOH that can be detected prevents this technique from being used to detect specific loci. Even using digital PCR, these methods cannot detect the relatively small amounts of LOH. Furthermore, this technique remains limited to studying known specific loci that occur in specific cancer types. Therefore, this method cannot be used, or cannot be effectively used, as a general means of screening for various cancers.

[0006] Besides limitations in screening for the presence or absence of cancer, current technologies are even less effective at providing prognostic diagnoses and monitoring treatment outcomes for cancer patients (such as recovery after surgery, chemotherapy, immunotherapy, or targeted therapy). This is because such technologies are typically expensive (e.g., imaging techniques), have low accuracy, low efficiency, and low sensitivity, and if imaging techniques are used, patients may be exposed to radiation.

[0007] Therefore, the ideal invention is a new technology that can effectively provide screening, prognosis, and monitoring for cancer patients. Invention Overview

[0008] This invention provides systems, instruments, and methods with multiple embodiments for detecting cancer-related genetic abnormalities. Analysis of biological samples, including cell-free DNA fragments, identifies imbalances in chromosomal regions within tumors, such as those resulting from chromosomal fragment deletions and / or amplifications. Utilizing chromosomal regions with multiple loci can yield higher efficiency and / or accuracy. These imbalances can then be used for the diagnosis or screening of potential cancer patients and for prognosis. The severity and number of imbalances can be used to provide cancer diagnosis, screening, prognosis, and monitoring. Furthermore, patients can be tested at different time points to track the degree and number of imbalances in each of one or more chromosomal regions, enabling cancer screening and prognosis, as well as monitoring (e.g., post-treatment).

[0009] According to one embodiment, a method is provided for analyzing biological samples of an organism for deletions or amplifications of a cancer-related chromosome. The biological sample includes nucleic acid molecules derived from normal cells and potentially cancer-related cells. At least some nucleic acid molecules in the sample are free. First and second haplotypes are determined for normal cells of the organism in a first chromosomal region. The first chromosomal region includes a first plurality of heterozygous loci. Each of the plurality of nucleic acid molecules in the sample has identifiable location information and its respective allele information in a reference genome of the identified organism. The location and the determined alleles are used to determine a first group of nucleic acid molecules derived from the first haplotype and a second group of nucleic acid molecules derived from the second haplotype. A first value corresponding to the first group of nucleic acids and a second value corresponding to the second group of nucleic acids are calculated using a computer system. Each value defines a property of each group of nucleic acid molecules (e.g., the average size or number of molecules in each group). The first value is compared with the second value to determine whether the first chromosomal region exhibits a deletion or amplification in any cancer-related cells.

[0010] According to another embodiment, a method for analyzing biological samples of an organism is provided. The biological sample includes nucleic acid molecules derived from normal cells and potentially cancer-related cells. At least some nucleic acid molecules in the sample are free. Multiple non-overlapping chromosomal regions of the organism are identified. Each chromosomal region includes multiple loci. Each of the multiple nucleic acid molecules in the sample has corresponding positional information in a reference genome of the identified organism. For each chromosomal region, the chromosomal region corresponding to each group of nucleic acid molecules is identified based on the identified position. Each group of nucleic acid molecules includes at least one nucleic acid molecule falling on each of the multiple loci within the chromosomal region. Values ​​in each group are calculated using a computer system, where each value defines the properties of the nucleic acid molecules in each group. The values ​​are compared with reference values ​​to determine whether chromosomal regions have deletions or amplifications. Chromosomal regions classified as having chromosomal deletions or amplifications are then quantitatively determined.

[0011] According to another embodiment, a method is provided for determining the progression of chromosomal aberrations in an organism using biological samples, wherein the biological samples include nucleic acid molecules derived from normal cells and potentially cancer-related cells. In the biological samples, at least some nucleic acid molecules are free. One or more non-overlapping chromosomal regions are identified against a reference genome of the organism. Each chromosomal region includes multiple loci. Samples obtained from the organism at different time points are analyzed to determine the progression of the disease. For each sample, each of the multiple nucleic acid molecules in the sample has a location in the reference genome of the identified organism. For each chromosomal region, the chromosomal region corresponding to each set of nucleic acid molecules is identified based on the identified location. Each set of nucleic acid molecules includes at least one nucleic acid molecule at each of the multiple loci located in the chromosomal region. A computer system calculates values ​​for each set of nucleic acid molecules. Each value defines the properties of each set of nucleic acid molecules. The values ​​are compared with reference values ​​to determine whether deletions or amplifications exist in a first chromosomal region. Then, the progression of chromosomal aberrations in the organism is determined at multiple time points using the possible deletions or amplifications in each chromosomal region.

[0012] Other embodiments of the invention are directed to systems, portable devices, and computer-readable media relating to the methods described herein.

[0013] The features and advantages of the present invention can be better understood by referring to the following detailed description of the invention and the accompanying drawings. Attached Figure Description

[0014] Figure 1 It explains that there are deleted or aberrant chromosomal regions in cancer cells.

[0015] Figure 2 It describes the presence of amplified and aberrant chromosomal regions in cancer cells.

[0016] Figure 3 Table 300 shows the different types of cancer, the associated regions, and their corresponding abnormalities.

[0017] Figure 4 According to an embodiment of the present invention, the dosage corresponding to the chromosomal region that does not show aberrations in cancer cells is measured in plasma.

[0018] Figure 5 According to an embodiment of the present invention, a dose corresponding to a missing chromosomal region 510 in cancer cells is measured in plasma to determine the corresponding missing region.

[0019] Figure 6 According to an embodiment of the present invention, a dose corresponding to an amplified chromosomal region 610 present in cancer cells is measured in plasma to determine the amplified region.

[0020] Figure 7 The present invention describes an RHDO analysis of plasma DNA from patients with hepatocellular carcinoma (HCC) targeting a chromosomal segment located on chromosome 1p, wherein the chromosomal segment shows the presence of monoallelic amplification in tumor tissue.

[0021] Figure 8 The present invention describes the changes in the length distribution of nucleic acid fragments in the plasma corresponding to the two haplotypes of the chromosomal region when a tumor has chromosomal deletions.

[0022] Figure 9 The present invention clarifies the changes in the length distribution of nucleic acid fragments in the plasma corresponding to the two haplotypes of the chromosomal region of a tumor when chromosomal amplification is present.

[0023] Figure 10 A flowchart illustrating a method for analyzing haplotypes of biological samples of an organism to determine whether chromosomal regions exhibit deletions or amplifications, according to an embodiment of the present invention, is provided.

[0024] Figure 11 The invention describes an embodiment of the invention, in which a missing region 1110 and a subregion 1130 exist in cancer cells, and a quantitative measurement of the region in plasma is performed to determine the missing region.

[0025] Figure 12 This paper describes how, according to an embodiment of the present invention, RHDO analysis can be used to plot the location of distortion.

[0026] Figure 13 The classification of RHDOs starting from another direction is illustrated according to an embodiment of the present invention.

[0027] Figure 14 A flowchart of a method 1400 for analyzing biological samples of an organism using multiple chromosome regions, according to an embodiment of the present invention.

[0028] Figure 15 According to an embodiment of the invention, Table 1500 illustrates the sequencing depth required for different numbers of aberrant fragments at different relative percentage concentrations of tumor nucleic acids. Figure 15 It provides an assessment of the number of analyte molecules at different relative percentage concentrations of tumor nucleic acids in a sample.

[0029] Figure 16The principle of measuring the relative percentage concentration of tumor nucleic acids in plasma by relative haplotype dose (RHDO) analysis according to an embodiment of the present invention is explained. HapI and HapII represent two haplotypes in non-tumor tissues according to an embodiment of the present invention.

[0030] Figure 17 According to an embodiment of the present invention, a flowchart is provided for a method of determining the process of chromosomal aberrations in an organism using a biological sample containing nucleic acid molecules.

[0031] Figure 18A SPRT curves for RHDO analysis of chromosomal segments on the q-arm of chromosome 4 in patients with cancer are shown. Dots represent the ratio of cumulative counts at all analyzed loci upstream of each heterozygous locus. Figure 18B SPRT curves for RHDO analysis of chromosomal segments on the q-arm of chromosome 4 in post-treatment patients are shown.

[0032] Figure 19 It shows common chromosomal aberrations found in HCC.

[0033] Figure 20A The results show the proportion of normalized sequencing reads obtained by analyzing HCC patients and healthy controls using targeted analysis (e.g., targeted enrichment capture sequencing technology). Figure 20B The results show the analysis of nucleic acid length in plasma obtained after targeted enrichment capture sequencing in 3 HCC patients and 4 healthy controls.

[0034] Figure 21 The diagram shows a Circos plot of an HCC patient obtained according to an embodiment of the present invention, which depicts data obtained from sequencing tag counting of plasma DNA.

[0035] Figure 22 This demonstrates a sequencing tag counting analysis of plasma samples from chronic hepatitis B virus (HBV) carriers who do not have HCC, according to an embodiment of the invention.

[0036] Figure 23 This demonstrates a sequencing tag counting analysis performed on plasma samples from patients with stage III nasopharyngeal carcinoma (NPC) according to an embodiment of the invention.

[0037] Figure 24 This demonstrates a sequencing tag counting analysis performed on plasma samples from patients with stage IV NPC according to an embodiment of the present invention.

[0038] Figure 25The diagram shows the distribution of cumulative plasma DNA length frequency in regions of tumor tissue with loss of heterozygosity (LOH) according to an embodiment of the invention.

[0039] Figure 26 The relationship between ΔQ and the length of plasma DNA in the LOH region was shown. According to an embodiment of the invention, ΔQ reaches 0.2 when the length is 130 bp.

[0040] Figure 27 The diagram shows the cumulative frequency distribution of plasma DNA length in regions of tumor tissue where chromosomal amplification exists, according to an embodiment of the present invention.

[0041] Figure 28 This demonstrates the relationship between ΔQ and plasma DNA length for the amplified region, according to an embodiment of the invention.

[0042] Figure 29 A block diagram of a computer system example 900 available according to an embodiment of the present invention is shown.

[0043] definition

[0044] As used herein, the term “biological sample” means any sample taken from a subject (e.g., a human, a person with cancer, a person suspected of having cancer, or another organism) and which includes one or more nucleic acid molecules of interest.

[0045] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in single-stranded or double-stranded form, and their polymers. Unless specifically defined, the term covers nucleic acids including known analogs of natural nucleotides that have similar binding properties to a reference nucleic acid and are metabolized in the same manner as naturally formed nucleotides. Unless otherwise stated, a particular nucleic acid sequence also implies the inclusion of variants with conserved modifications (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), copy number variants, complementary sequences, and explicitly specified sequences. Specifically, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is replaced by a mixture of bases and / or deoxyinosine residues (Batzer). et al ., Nucleic Acid Res . 19:5081 (1991); Ohtsuka et al ., J. Biol. Chem 260:2605-2608 (1985); and Rossolini et al ., Mol. Cell. Probes8:91-98 (1994)). The term nucleic acid covers, but is not limited to: genes, complementary DNA (cDNA), messenger RNA (mRNA), small non-coding RNA, microRNA (miRNA), Piwi-interacting RNA, and short hairpin RNA (shRNA) or other sequences on chromosomes encoded by genes or loci.

[0046] The term "gene" refers to a segment of DNA associated with the production of polypeptide chains or transcribed RNA products. It may include regions before and after the coding region (leader and tail regions), as well as spacer sequences (introns) between individual coding segments (exons).

[0047] As used herein, the terms "clinically relevant nucleic acid sequence" or "clinically relevant chromosomal region" (or region / segment to be tested) can refer to a large genomic sequence fragment with a potential imbalance to be tested, or a polynucleotide sequence corresponding to a large genomic sequence of its own. Examples include deleted or amplified, or potentially deleted or amplified, genomic segments (including simple duplications), or large regions including subregions of that segment. In some implementations, multiple clinically relevant nucleic acid sequences, or multiple equivalent markers of clinically relevant nucleic acid sequences, can provide data for detecting imbalances in the region. For example, data obtained from five non-contiguous sequences on a chromosome can be summed to determine a potential imbalance, thereby effectively reducing the required sample dose to 1 / 5.

[0048] As used herein, the terms "reference nucleic acid sequence" or "reference chromosomal region" refer to a nucleic acid sequence whose dose distribution or length distribution is compared to that of the region to be tested. Examples of reference nucleic acid sequences include chromosomal regions that do not contain deletions or amplifications, complete genomes (e.g., normalized by total sequencing tag counts), regions obtained from one or more known normal samples (which may be the same region as the sample to be tested), or chromosomal regions of a specific haplotype. Such reference nucleic acid sequences may be present endogenously in the sample or added exogenously during sample processing or analysis. In some embodiments, the reference chromosomal region demonstrates a length distribution representing a disease-free healthy state. In other embodiments, the reference chromosomal region demonstrates a quantitative profile representing a disease-free healthy state.

[0049] As used herein, the term "based on" means "at least partially based on" and refers to a value (or result) used in the determination of another value, such as in the relationship between the inputs and outputs of a method. As used herein, the term "derive" also refers to the relationship between the inputs and outputs of a method, such as in the calculation of a formula.

[0050] As used herein, the term "parameter" refers to a numerical value that characterizes the quantitative relationship between sets of quantitative data and / or between sets of quantitative data. For example, the ratio (or a function of the ratio) of a first quantity of a first nucleic acid sequence to a second quantity of a second nucleic acid sequence is a parameter.

[0051] As used in this article, the term "locus" is the location or address of a nucleotide (or base pair) of any length that can vary in the genome.

[0052] As used herein, the term “sequence imbalance” or “abnormality” refers to any significant deviation of the quantity of a clinically relevant chromosomal region from at least one defined threshold of a reference quantity. Sequence imbalance can include chromosomal dosing imbalance, allele imbalance, mutation dosing imbalance, copy number imbalance, haplotype dosing imbalance, and other similar imbalances. For example, allele imbalance can occur because a tumor has an allele of a deleted gene or an allele of an amplified gene, or differential amplification of two alleles in its genome, thereby creating an imbalance at a specific locus in the sample. As another example, a patient may have a genetic mutation in a tumor suppressor gene. The patient can then continue to develop a tumor in which the non-mutated allele of the tumor suppressor gene is deleted. Therefore, a mutation dosing imbalance exists in the tumor. When a tumor releases its DNA into a patient's plasma, the tumor DNA mixes with the patient's normal somatic structural DNA in the plasma. The mutation dosing imbalance present in this DNA mixture can be detected using the methods described herein.

[0053] As used herein, the term "haplotype" refers to a combination of alleles at multiple loci, wherein these alleles are transmitted together to offspring on the same chromosome or chromosomal region. A haplotype can refer to as few as one pair of loci or chromosomal regions, or an entire chromosome. The term "allele" refers to alternative DNA sequences at the same physical location on the same chromosome that may yield the same or different phenotypic traits. In any given diploid organism, each chromosome has two copies (except for the sex chromosomes in male subjects), and the genotype of each gene includes the paired alleles present at that locus; if they are the same, the individual is homozygous, and if they are different, the individual is heterozygous. An organismal population or species typically includes multiple alleles at various loci in different individuals. Loci where more than one allele is found in a population are called polymorphic loci. The degree of variation in alleles at a locus can be measured by the number of alleles (i.e., the degree of polymorphism) or the proportion of heterozygotes in the population (i.e., the proportion of heterozygotes). As used in this article, the term "polymorphism" refers to any variation in the human genome between individuals, regardless of its frequency. Examples of such variations include, but are not limited to, single nucleotide polymorphisms, simple tandem repeat polymorphisms, insertion-deletion polymorphisms, mutations (which may be the cause of disease), and copy number variations.

[0054] The term "sequencing tag" refers to the sequence determined from all or part of a nucleic acid molecule (such as a fragment of DNA). Typically, only one end of the fragment is sequenced, for example, about 30 bp. The sequencing tag is then aligned to a reference genome. Alternatively, both ends of the fragment can be sequenced, generating two sequencing tags, which can provide higher alignment accuracy and also provide fragment length information.

[0055] The term "universal sequencing" refers to sequencing fragments by ligating known adapter sequences to the ends of the fragments, with primers used for sequencing complementarily pairing to the adapter sequences. Therefore, any fragment can be sequenced using the same primers, and thus sequencing can be random.

[0056] The term "size distribution" refers to any value or set of other measures representing the length, mass, weight, or size of a molecule corresponding to a particular group (e.g., a fragment obtained from a particular haplotype or a particular chromosomal region). Various implementations may use various size distributions. In some implementations, the size distribution relates to the relative order of lengths of one chromosome-related fragment relative to other chromosome-related fragments (e.g., mean, median, or geometric mean). In other implementations, the size distribution may relate to a statistical value of the actual size of the chromosome fragments. In one implementation, the statistical value may include any mean, geometric mean, or median of the chromosome fragments. In another implementation, the statistical value may include the total length of fragments below a certain threshold, which may be divided by the total length of all fragments or at least fragments below a certain larger truncation value.

[0057] As used herein, the term “classification” refers to any number or other characteristic associated with a particular property of a sample. For example, a “+” sign (or the word “positive”) can indicate that a sample is classified as having a chromosomal aberration, whether deletion or amplification. Classification can be binary (e.g., positive or negative) or have higher levels of classification (e.g., levels of 1 to 10 or 0 to 1). The terms “truncation” and “threshold” refer to predetermined values ​​used in the operation. For example, a truncation size can refer to a value that corresponds to segments that will be excluded. A threshold can be a value above or below which a particular classification can be applied. Any of these terms may be used in any of the contexts herein.

[0058] The term "level of cancer" can refer to the presence of cancer, the stage of cancer, the size of the tumor, the extent of deletion or amplification involving chromosomal regions (e.g., double or triple duplication amplification), and / or other measures of cancer severity. The level of cancer can be quantitative or other characteristics. The level can be zero. Furthermore, the level of cancer also includes pre-malignant or precancerous states associated with deletions or amplifications. Invention Details

[0059] Cancerous tissue (tumors) can exhibit aberrations, such as deletions or amplifications of chromosomal regions. Tumors can release DNA fragments into the body's fluid tissues, such as peripheral blood. Multiple approaches can identify aberrations, and thus tumors, by analyzing DNA fragments relative to normal values ​​(expected values) of DNA in chromosomal regions.

[0060] The exact size and location of deletions or amplifications can change. A specific chromosomal region may contain known aberrations common to cancer or a particular cancer (thus enabling diagnosis of that specific cancer). When a specific region is unknown, systematic methods used to analyze the entire genome or a large portion of the genome can be used to detect aberrant regions, which may be scattered throughout the genome and whose size (e.g., the number of deleted or amplified bases) varies. Chromosomal regions can be tracked at different time points to detect changes in the severity of an aberration or changes in several aberrant regions. This tracking can provide crucial information for cancer screening, prognostic diagnosis, and tumor monitoring (e.g., after treatment or to detect recurrence or tumor progression).

[0061] This invention begins with an example of chromosomal aberrations in cancer. It then discusses examples of methods for detecting chromosomal aberrations by detecting and analyzing cell-free DNA in biological samples. Once a method for detecting aberrations in a single chromosomal region is established, methods for systematically applying the detection of aberrations in many chromosomal regions to the screening (diagnosis) and prognosis of cancer patients are further described. Furthermore, this invention details methods for detecting one or more regions at multiple time points to track cancer-related quantitative indicators obtained from examining chromosomal aberrations, thereby providing screening, prognosis, and monitoring for patients. Examples are then discussed.

[0062] I. Examples of chromosomal aberrations in cancer

[0063] Chromosomal aberrations are prevalent in cancer cells. Furthermore, specific cancers exhibit characteristic patterns of chromosomal aberrations. For example, increased amounts of DNA at chromosome arms 1p, 1q, 7q, 15q, 16p, 17q, and 20q, and decreased amounts of DNA at 3p, 4q, 9p, and 11q, are commonly detected in hepatocellular carcinoma (HCC). Previous studies have demonstrated that such genetic aberrations can also be detected in peripheral blood DNA from cancer patients. For instance, in peripheral blood DNA molecules from patients with lung and head and neck cancer, studies have indicated the detection of loss of heterozygosity (LOH) at specific loci (Chen XQ, et al. Nat Med 1996; 2: 1033-5; Nawroz H, et al. Nat Med 1996; 2: 1035-7). Genetic variations detected in plasma or serum are consistent with those found in tumor tissue. However, since tumor-derived DNA constitutes only a small portion of the total cell-free DNA in peripheral blood, the allelic imbalance caused by LOH in tumor cells is usually weak. Many researchers have developed digital polymerase chain reaction (PCR) technology (Vogelstein B, Kinzler KW. Proc Natl Acad Sci US A. 1999; 96: 9236-41; Zhou W, et al. Nat Biotechnol 2001; 19: 78-81; Zhou W, et al. Lancet. 2002;359: 219-25) to accurately quantify different alleles at loci in peripheral blood DNA molecules (Chang HW, et al. J Natl Cancer Inst. 2002; 94: 1697-703). Digital PCR is more sensitive than real-time PCR or other DNA quantification methods used to detect weak allelic imbalances caused by LOH at specific loci in tumor DNA. However, digital PCR still has difficulties in identifying extremely weak allelic imbalances at specific loci; therefore, the implementation method described in this paper analyzes different regions of the chromosome in a comprehensive manner.

[0064] Furthermore, the techniques described herein can be applied to detect pre-malignant or precancerous conditions. Examples of such conditions include cirrhosis and cervical intraepithelial neoplasia (CIN). Cirrhosis is a pre-hepatocellular carcinoma (HCC) condition, while CIN is a pre-cervical cancer condition. These pre-malignant conditions have reportedly undergone various molecular alterations during their development into malignant tumors. For example, the presence of LOH at chromosome arms 1p, 4q, 13q, and 18q, and the simultaneous deletion at more than three loci, are associated with an increased risk of HCC in patients with cirrhosis (Roncalli Met al. Hepatology 2000;31:846-50). In addition, these pre-malignant lesions release DNA into the peripheral blood circulation, albeit at low concentrations. The techniques described can detect deletions or amplifications by analyzing DNA fragments in plasma and measuring the concentration (including relative percentage concentration) of pre-malignant cell-free DNA in the plasma. Such aberrations can be easily detected (e.g., by sequencing depth or the number of such changes detected), and their concentration can predict the likelihood or speed at which a cancer condition will worsen into a full-blown outbreak.

[0065] A. Deletion of chromosomal regions

[0066] Figure 1 This study describes the presence of deleted or aberrant chromosomal regions in cancer cells. Normal cells show two haplotypes, Hap I and Hap II. (Example...) Figure 1 As shown, in each of the multiple heterozygous loci 110 (also known as single nucleotide polymorphisms, SNPs), both Hap I and Hap II have sequence information. In cancer-associated cells, Hap II has a missing chromosomal region 120. For example, cancer-associated cells may originate from tumors (e.g., malignant tumors), from metastatic lesions of tumors (e.g., in local lymph nodes or in distant organs), or from precancerous or pre-malignant lesions, such as those described above.

[0067] In chromosomal region 120 of cancer cells, one of the two homologous haplotypes is deleted. Because of the absence of other corresponding alleles on the deleted homologous chromosome, all heterozygous SNPs 110 appear homozygous. Therefore, this type of chromosomal aberration is called loss of heterozygosity (LOH). In region 120, the non-deleted alleles of these SNPs represent one of the two haplotypes that can be found in normal tissue. Figure 1In the example shown, haplotype I (Hap I) at region 120 of the LOH can be determined using the genotype of the tumor tissue. Other haplotypes (Hap II) can be determined by comparing the genotype of normal tissue with the genotype on the surface of cancerous tissue. Hap II can be constructed by linking all the missing alleles. That is, alleles present only at region 120 in normal cells (but missing at region 120 in cancer cells) are determined to be the same haplotype, i.e., Hap I. Through this analysis, the haplotypes of patients corresponding to all chromosomal regions in tumor tissue exhibiting LOH can be determined (e.g., hepatocellular carcinoma (HCC) patients). This method is only useful for cancer cell analysis and is only effective for determining haplotypes in region 120, but it elucidates the missing regions in chromosomes very well.

[0068] B. Amplification of chromosomal regions

[0069] Figure 2 This study describes how cancer cells exhibit amplified and aberrant chromosomal regions. Normal cells show two haplotypes, Hap I and Hap II. (Example...) Figure 2 As shown, in each of the multiple heterozygous loci 210, both Hap I and Hap II have sequence information. In tumor cells, Hap II has a 2-fold amplified (replicated) chromosomal region 220.

[0070] Similarly, for regions of tumor tissue with monoallelic amplification, methods such as microarray analysis can be used to detect amplified alleles at SNP 210. One of two haplotypes can be determined by linking all amplified alleles in chromosomal region 220 together (e.g., ...). Figure 2 (Hap II in the example shown). Amplified alleles at a specific locus can be determined by comparing the number of alleles at each locus. Then, other haplotypes (Hap I) can be determined by linking the non-amplified alleles together. This method is only useful for cancer cell analysis and is only effective for determining haplotypes in region 220, but it effectively elucidates amplified regions in the chromosome.

[0071] Amplification can originate from regions of more than two chromosomes, or from a repetition of a gene on a single chromosome. A region can be tandemly copied, or a region can be a small chromosome comprising one or more copies of that region. Furthermore, amplification can also originate from a gene on a copied chromosome, with the copy product inserted into a different chromosome or a different region of the same chromosome. This type of insertion constitutes a type of amplification.

[0072] II. Selection of Chromosomal Regions

[0073] Genomic aberrations in cancerous tissue can be detected in samples such as plasma and serum when the cancerous tissue contributes at least a portion of cell-free DNA (and potentially intracellular DNA). The challenge in detecting these aberrations is that tumors or cancers can be quite small, resulting in relatively weak DNA contribution from cancer cells. Consequently, the amount of cell-free DNA with aberrations is relatively small, making detection extremely difficult. At a single locus in the genome containing the aberration to be detected, there may not be sufficient DNA. The method described herein overcomes this difficulty by analyzing DNA at chromosomal regions (haplotypes) comprising multiple loci, thereby clustering small changes at individual loci into perceptible differences based on haplotype. Thus, analyzing multiple loci within a region provides greater accuracy and precision, and reduces false positives and false negatives.

[0074] Furthermore, the aberrant regions can be quite small, making them difficult to identify. If only one locus or specific loci are used, aberrations not located at those loci will be missed. As described in this paper, methods can be used to study entire regions, thereby discovering aberrations present in subregions within those regions. When the analyzed region covers the entire genome, the entire genome can be analyzed to discover aberrations of varying lengths and locations, as described in more detail below.

[0075] To illustrate these points, as shown above, regions can exhibit distortion. However, the region used for analysis must be carefully selected. The length and location of the region can alter the results and thus affect the analysis. For example, if the analysis... Figure 1 In the first region shown, no aberrations were detected. If the second region is analyzed, for example using the method described herein, aberrations can be detected. When analyzing a larger region that includes both the first and second regions, one encounters the problem that only a portion of the larger region may have aberrations, making it more difficult to identify any aberrations, and one faces the problem of determining the exact location and length of the aberration. Various embodiments of the invention can solve some and / or all of these problems. The description for selecting regions also applies to haplotypes using the same chromosomal region or haplotypes using two different chromosomal regions.

[0076] A. Selecting specific chromosomal regions

[0077] In one implementation, specific regions can be selected based on knowledge of the cancer or the patient. For example, the region may be a known abnormality prevalent in many cancers or a specific type of cancer. The exact length and location of the corresponding region can be obtained by referring to well-known literature related to the type of cancer or specific risk factors present in the patient. Furthermore, the patient's tumor tissue can be obtained and analyzed to identify the abnormal region, as described above. Currently, such techniques require obtaining cancer cells (which may be impractical for newly diagnosed patients), but they can be used to identify regions monitored at different time points in the same patient (e.g., after surgical removal of cancerous tissue, or after chemotherapy, immunotherapy, or targeted therapy, or to detect tumor recurrence or progression).

[0078] People can identify more than one specific region. Each of these regions can be analyzed independently, or different regions can be analyzed together. Furthermore, these regions can be subdivided, thereby providing greater accuracy in locating distortions.

[0079] Figure 3 Table 300 shows different types of cancer, associated regions, and their corresponding aberrations. Column 310 lists the different cancer types. The implementation described herein can be used for any type of cancer associated with the aberration; therefore, this list is merely illustrative. Column 320 shows multiple regions (e.g., large regions such as 7p or more specific regions such as 17q25) where an increase (amplification) is associated with a specific cancer in the same row. Column 330 shows regions where a deletion (removal) can be found. Column 340 lists references that discuss the relevance of these regions to specific cancers.

[0080] According to the methods described herein, these regions with potential chromosomal aberrations can be considered as candidate chromosomal regions for aberration analysis. Examples of other genomic regions altered in cancer can be found in the Cancer Genome Anatomy Project (cgap.nci.nih.gov / Chromosomes / RecurrentAberrataions) and the Atlas of Genetics and Cytogenetics in Oncology and Haematology (atlasgeneticsoncology.org Tumors / Tumorliste.html) databases.

[0081] As can be seen, the identified areas can be quite large, while others can be more specific. The aberration does not necessarily encompass the entire area identified in the table. Therefore, such clues for this type of aberration cannot be precisely located for a specific patient, but can more often serve as a rough guide for the large areas to be analyzed. Such large areas can include many sub-regions (which can be of equal size). They can be used for individual analyses as well as combined analyses (details of which are described herein). Thus, based on the specific circumstances of the cancer to be examined, multiple implementations can combine various aspects of selecting large areas, but more general techniques, as described below, can also be used.

[0082] B. Select any chromosome region

[0083] In another implementation, the chromosomal region to be analyzed is arbitrarily selected. For example, the genome can be divided into regions of 1 megabase (Mb) in length, or other predetermined segments of length, such as 500 kb or 2 Mb. If a region is 1 Mb, there are approximately 3,000 regions in the human genome, since there are approximately 3 billion bases in the haploid human genome. Each of these regions can be analyzed, as discussed in more detail below.

[0084] The determination of such regions is not based on any knowledge of the cancer or the patient, but rather on the systematic segmentation of the genome into regions to be analyzed. In one implementation, when a chromosome does not have the length of several predetermined segments (e.g., not divisible by 1 million bases), the final region of the chromosome can be smaller than the predetermined length (e.g., less than 1 Mb). In another implementation, each chromosome can be divided into regions of equal length (or approximately equal—within rounding error) based on the total length of the chromosome and the number of segments to be created (which typically varies within a chromosome). In such implementations, the length of segments on each chromosome may differ.

[0085] As mentioned above, specific regions can be identified based on the specific cancer being examined, but these regions can also be subdivided into smaller regions (e.g., sub-regions of equal size covering a larger specific region). In this way, aberrations can be precisely located. In the discussion below, any general reference region for a chromosomal region can be a specifically identified region and / or an arbitrarily selected region.

[0086] III. Detection of aberrations in specific haplotypes

[0087] This section describes a method for detecting aberrations in a single chromosomal region by analyzing a biological sample containing cell-free DNA. In the embodiments described in this section, a single chromosomal region is a region containing multiple heterozygous (different allele) loci, thereby allowing differentiation of two haplotypes by identifying the specific allele at a given locus. Therefore, a given nucleic acid molecule (e.g., a fragment of cell-free DNA) can be identified as belonging to a specific haplotype from two haplotypes. For example, the fragment can be sequenced to obtain a sequence tag aligned to the chromosomal region, and the haplotype at the heterozygous locus to which the allele belongs can then be identified. Two common techniques for determining aberrations in a specific haplotype (Hap) are described below: tag counting and size analysis.

[0088] A. Haplotype determination

[0089] To distinguish between the two haplotypes, the two haplotypes in the chromosomal region must first be determined. For example, it can be determined that... Figure 1 The normal cells show two haplotypes, Hap I and Hap II. Figure 1 In this study, the haplotype includes a first plurality of loci 110, which is heterozygous and allows for the differentiation of two haplotypes. This first plurality of loci covers the chromosomal region to be analyzed. Alleles at different heterozygous loci (hets) can be determined first, and then the patient's haplotype can be determined.

[0090] SNP allele haplotypes can be determined using single-molecule analysis methods. Examples of such methods have been described by researchers Fan et al. (Nat Biotechnol. 2011;29:51-7), Yang et al. (Proc Natl Acad Sci US A. 2011;108:12-7), and Kitzman et al. (Nat Biotechnol. 2011 Jan;29:59-63). Furthermore, an individual's haplotype can be determined by analyzing the genotypes of family members (e.g., parents, siblings, and children). Examples include the methods described by Roach et al. (Am J Hum Genet. 2011;89(3):382-97) and Lo et al. (SciTransl Med. 2010;2:61ra91). In another embodiment, an individual's haplotype can be determined by comparing the genotyping results of tumor tissue with the genotyping results of normal structural genomes. The genotypes of these subjects can be obtained through microarray analysis (e.g., using t).

[0091] In addition, haplotypes can also be constructed using other methods familiar to those skilled in the art. Examples of such methods include haplotype determination based on single-molecule analysis, such as digital PCR (Ding C and Cantor CR. ProcNatl Acad Sci USA 2003; 100: 7449-7453; Ruano G et al. Proc Natl Acad Sci USA 1990; 87: 6296-6300), chromosome selection or isolation (Yang H et al. Proc Natl Acad Sci US A 2011; 108: 12-17; Fan HC et al. Nat Biotechnol 2011; 29: 51-57), sperm haplotype analysis (Lien S et al. Curr Protoc Hum Genet 2002; Chapter 1: Unit 1.6), and imaging techniques (Xiao M et al. Hum Mutat 2007; 28: 913-921). Other methods include haplotype analysis techniques based on allele-specific PCR (Michalatos-Beloin S et al. Nucleic Acids Res 1996; 24: 4841-4843; Lo YMD et al. Nucleic Acids Res 19: 3561-3567), cloning, and restriction enzyme digestion (Smirnova AS et al. Immunogenetics 2007; 59:93-8). Other methods are based on the distribution of haplotype blocks and linkage disequilibrium in a population, allowing the haplotype of a subject to be derived through statistical evaluation (Clark AG. Mol Biol Evol 1990; 7:111-22; 10:13-9; Salem RM et al. Hum Genomics 2005;2:39-66).

[0092] Another method for determining haplotypes in regions of locus of leukemia (LOH) is through genotyping of normal and tumor tissues (if tumor tissue is available) from the subject. In the presence of LOH, tumor tissues with a very high relative percentage concentration of tumor cells will show epigenetic homozygosity at all SNP loci within the region where LOH is present. The genotypes of these SNP loci include one haplotype ( Figure 1The LOH region is shown as Hap I. On the other hand, in normal tissues, SNP loci within the LOH region show heterozygosity. Alleles present in normal tissues but not in tumor tissues include other haplotypes ( Figure 1 Hap II of the LOH region is shown.

[0093] B. Relative haplotype dose (RHDO) analysis

[0094] As mentioned above, chromosomal aberrations involving the amplification or deletion of one of the haplotypes in a chromosomal region can lead to a dosage imbalance of the two haplotypes in the corresponding chromosomal region within tumor tissue. In the plasma of individuals with tumor growth, a portion of peripheral blood DNA originates from tumor cells. Because tumor-derived DNA is present in the plasma of cancer patients, this imbalance will also be present in their plasma. The dosage imbalance of the two haplotypes can be detected by counting the number of molecules derived from each haplotype.

[0095] LOH chromosomal regions were observed in tumor tissue (e.g.) Figure 1 As shown in region 120, Hap I is relatively excessive in peripheral blood DNA molecules (fragments) compared to Hap II due to the lack of contribution from Hap II in tumor tissue. In chromosomal regions where copy number amplification is observed in tumor tissue, Hap II is relatively excessive compared to Hap I in regions affected by monoallelic amplification of Hap II due to the additional dose released by the tumor tissue. To determine whether an excess or deficiency is present, various methods can be used to measure certain DNA fragments derived from Hap I or Hap II in the sample, such as based on universal sequencing and alignment analysis, or using digital PCR and sequence-specific probes.

[0096] After sequencing multiple DNA fragments obtained from the plasma (or other biological samples) of cancer patients to generate sequence tags, sequencing tags corresponding to alleles on two haplotypes can be identified and counted. The number of sequencing tags corresponding to each of the two haplotypes is then compared to determine whether the two haplotypes are present equally in the plasma. In one embodiment, a sequential probability ratio test (SPRT) can be used to determine whether there is a significant difference in the presentation of the two haplotypes in the plasma. A statistically significant difference indicates the presence of chromosomal aberrations in the analyzed chromosomal region. Furthermore, the quantitative difference between the two haplotypes in the plasma can be used to estimate the relative percentage concentration of tumor-derived DNA in the plasma, as described below.

[0097] The diagnostic methods described in this application for determining the characteristics of DNA fragments (e.g., their location in the human genome) are not limited to using massively parallel sequencing as a detection platform according to embodiments of the invention. Furthermore, these diagnostic methods can also be used, for example, but not limited to, microfluidic digital PCR systems (e.g., fluidic digital array systems), droplet digital PCR systems (e.g., RainDance and QuantaLife), BEAM-ing systems (i.e., systems based on bead, emulsion PCR, amplification, and magnetism) (Diehl et al. ProcNatlAcadSci USA 2005; 102: 16368-16373), real-time PCR, mass spectrometry-based systems (e.g., SequenomMassArray systems), and multiplex linkage-dependent probe amplification (MLPA) analysis.

[0098] Normal area

[0099] Figure 4 This illustrates an embodiment of the invention, showing chromosomal regions within cancer cells that do not exhibit aberrations, and measurements taken in plasma. Chromosomal regions 410 can be selected by any method, such as methods based on the specific cancer to be tested or methods based on general screening methods (i.e., methods that use predetermined segments covering a large portion of the genome). To distinguish between two haplotypes, both haplotypes are first determined. Figure 4 Two haplotypes (Hap I and Hap II) of normal cells are shown with respect to chromosomal region 410. Each haplotype includes a first plurality of loci 420. These first plurality of loci 420 span the chromosomal region 410 to be analyzed. As shown, these loci are heterozygous in normal cells. Two haplotypes of cancer cells are also shown. In cancer cells, no regions were deleted or amplified.

[0100] also, Figure 4The number of allele counts at each haplotype with respect to each locus 420 is also shown. Furthermore, the cumulative totals for certain subregions of chromosome region 410 are provided. The number of allele counts corresponds to the number of DNA fragments (which correspond to a specific haplotype at each particular locus). For example, a DNA fragment containing locus 421 and having allele A is counted for Hap I. A DNA fragment having allele T is counted for Hap II. Where a fragment aligns to (i.e., whether it includes a specific locus) and which allele it includes can be determined in several ways, as mentioned herein. The count ratio at two haplotypes can be used to determine whether a statistically significant difference exists. This count ratio is also referred to herein as the odds ratio. Alternatively, the difference between two values ​​can be used; the difference can be normalized to the total number of fragments. Odds and differences (and their functions) are instances of parameters that are compared to a threshold to determine the classification of aberrations.

[0101] RHDO analysis can utilize all alleles (e.g., cumulative counts) on the same haplotype to determine the presence of any imbalance between two haplotypes in plasma, for example, in maternal plasma, as described in patent applications 12 / 940,992 and 12 / 940,993 of Lo, see above. This method can significantly increase the number of DNA molecules available for determining the presence of any imbalance, and can be used to distinguish cancer-induced imbalances from random fluctuations in non-cancer conditions (since allele counts in the absence of cancer or pre-existing conditions are randomly distributed, while cancer-induced imbalances are predetermined), thus achieving better statistical power. Compared to analyzing multiple SNP loci individually, the RHDO method can utilize the relative positions of alleles on two chromosomes (haplotype information), allowing for the joint analysis of alleles located on the same chromosome. In the absence of haplotype information, allele counts at different SNP loci cannot be summed together to statistically determine whether a haplotype is excessive or insufficient in plasma. Allele counting can be quantified by, but is not limited to, massively parallel sequencing (e.g., Illumina variant synthesis-side sequencing systems, Life Technologies' ligation-side sequencing (SOLiD) systems, Ion Torrent and Life Technologies' Ion Torrent sequencing systems, nanopore sequencing (nanoporetech.com), Roche 454 sequencing technology, digital PCR (e.g., microfluidic digital PCR (e.g., fluid (fluidigm.com))), BEAMing (bead, emulsion PCR, amplification, magnetic (inostics.com)), or droplet digital PCR (e.g., QuantaLife (quantalife.com) and RainDance (raindancetechnologies.com)) and real-time PCR). In other embodiments of the technology, high-throughput target capture sequencing (which utilizes liquid-phase capture (e.g., utilizing the Agilent SureSelect system, the Illumina TruSeq Custom Enrichment Kit (illumina.com / applications / sequencing / targeted_resequencing.ilmn), or MyGenostics GenCap Custom Enrichment system) can be used. (mygenostics.com / )) or array-based capture (e.g., using the Roche NimbleGen system).

[0102] exist Figure 4 In the example shown, a slight allele imbalance was observed at the first two SNP loci (24 and 26 for the first SNP, and 18 and 20 for the second SNP). However, the number of allele counts was statistically insufficient to determine whether a true allele imbalance existed. Therefore, allele counts on the same haplotype were added together until the cumulative allele counts for both haplotypes were sufficient to reach a statistically significant conclusion: there was no allele imbalance between the two haplotypes in chromosome region 410 (the fifth SNP in this example). After reaching a statistically significant classification, the cumulative count was reset (at the sixth SNP in this example). The cumulative count was then measured until the cumulative allele counts for both haplotypes were again sufficient to reach a statistically significant conclusion: there was no allele imbalance between the two haplotypes for a specific subregion of region 410. The total cumulative count could also be used for the entire region, but the previous method allowed different subregions to be tested, and unlike the entire region 410, the above test provided greater precision in determining the location of aberrations (i.e., subregions). Examples of statistical tests used to determine the presence of a true allele imbalance include, but are not limited to, the sequential probability ratio test (Zhou W, et al. Nat Biotechnol 2001; 19: 78-81; Zhou W, et al. Lancet. 2002; 359: 219-25), the t-test, and the chi-square test.

[0103] Detect and delete

[0104] Figure 5 The invention describes the deletion of chromosomal region 510 within cancer cells and measurements performed in plasma to determine the presence of the deletion. Figure 5 Two haplotypes (Hap I and Hap II) of chromosomal region 510 in normal cells are shown. Each haplotype includes a first plurality of heterozygous loci 520 that cover the chromosomal region 510 to be analyzed. Two haplotypes of cancer cells are also shown. In cancer cells, with respect to Hap II, region 510 has been deleted. Figure 4 similar, Figure 5 The number of allele counts for each locus 520 is also shown. In addition, the cumulative total is recorded for certain subregions within chromosome region 510.

[0105] Since tumor tissue typically comprises a mixture of tumor and non-tumor cells, a locus of absence (LOH) can be confirmed by the relative shift in the proportion of the two alleles at the locus within region 510. In this case, the deleted haplotype Hap II in region 510 can be determined by combining loci 520, which, compared to the corresponding loci in normal tissue, shows a relative reduction in the amount of DNA fragment. Haplotype fragments that appear more frequently are Hap I, which is retained in tumor cells. In some embodiments, it is ideal to perform a process that enriches the proportion of tumor cells in the tumor sample, thereby allowing for easier determination of both deleted and retained haplotypes. An example of such a process is microdissection (either manually or via laser capture technology).

[0106] Theoretically, for chromosomal regions exhibiting LOH in tumor tissue, each allele on Hap I is in relatively excess form in peripheral blood DNA, and the degree of allele imbalance depends on the relative percentage concentration of tumor DNA in the plasma. However, the relative amounts of the two alleles in the peripheral blood DNA sample simultaneously follow a Poisson distribution. Statistical analysis can be used to determine whether the observed allele imbalance is due to the actual presence of LOH in cancer tissue or due to random fluctuations. The ability to detect LOH-related allele imbalances in cancer depends on the number of peripheral blood DNA molecules analyzed and the relative percentage concentration of tumor DNA. Higher relative percentage concentrations of tumor DNA and a greater number of molecules used for analysis result in higher sensitivity and specificity for detecting allele imbalances.

[0107] exist Figure 5 In the example shown, a slight allele imbalance was observed at the first two SNP loci (24 and 22 for the first SNP, and 18 and 15 for the second SNP). However, the number of allele counts was statistically insufficient to determine whether a true allele imbalance existed. Therefore, allele counts on the same haplotype were added together until the cumulative allele counts for both haplotypes were sufficient to reach a statistical conclusion that an allele imbalance existed between the two haplotypes in region 510 (the fifth SNP in this example). In some embodiments, only the imbalance is known, and the specific type (deletion or amplification) is not determined. The cumulative count is then determined until the cumulative allele counts for both haplotypes are again sufficient to reach a statistical conclusion that an allele imbalance exists between the two haplotypes for a specific subregion of region 510. The total cumulative count can also be used for the entire region, which can be implemented using any of the methods described herein.

[0108] Detection of amplification of chromosomal regions

[0109] Figure 6 This paper describes the amplification of chromosomal region 610 within cancer cells and the measurement performed in plasma according to an embodiment of the present invention, thereby determining the amplified region. Besides LOH, amplification of chromosomal regions is frequently observed in cancer tissue. Figure 6 In the example shown, Hap II in chromosomal region 610 amplified to 3 copies in cancer cells. As shown, region 610 includes only 6 heterozygous loci, unlike the longer regions shown in previous figures. Amplification was identified as statistically significant at the 6 loci, with overpresentation being determined to be statistically significant. In some embodiments, only the imbalance is known, and the specific type (deletion or amplification) is not determined. In other embodiments, cancer cells can be obtained and analyzed. Such analysis can provide information about whether the imbalance is due to deletion (where the cancer cells are homozygous in relation to deleted regions) or amplification (where the cancer cells are heterozygous in relation to amplified regions). In other embodiments, the method described in part IV can be used to determine the presence of deletion or amplification, thereby analyzing the complete region (i.e., not analyzing individual haplotypes). If a region is overpresented, the aberration is amplification; if a region is underpresented, the aberration is deletion. Furthermore, region 620 was analyzed, and cumulative counts proved that no imbalance was present.

[0110] SPRT analysis for plasma RHDO analysis

[0111] For any chromosomal region containing a heterozygous locus, RHDO analysis can be used to determine the presence of any dose imbalance between two haplotypes in plasma. The presence of haplotype dose imbalance in plasma within these regions indicates the presence of tumor-derived DNA in the plasma sample. In one implementation, SPRT analysis can be used to determine whether the difference in the number of sequencing reads for Hap I and Hap II is statistically significant. In an example of SPRT analysis, we first determine the number of sequencing reads corresponding to each haplotype. Then, we determine a parameter (e.g., percentage concentration) that represents the proportion of sequencing reads contributed by potentially overpresented haplotypes (e.g., the proportion of reads for one haplotype divided by the proportion of reads for other haplotypes). In the LOH protocol, the potentially overpresented haplotype is the non-deleted haplotype, while in the protocol for monoallelic amplification of chromosomal regions, the potentially overpresented haplotype is the amplified haplotype. Next, this proportion is compared to two thresholds (an upper threshold and a lower threshold), constructed based on the null hypothesis (i.e., no haplotype dose imbalance exists) and the alternative hypothesis (i.e., a haplotype dose imbalance exists). If the proportion is greater than the upper threshold, it indicates a statistically significant imbalance between the two haplotypes in plasma. If the proportion is less than the lower threshold, it indicates no statistically significant imbalance between the two haplotypes. If the proportion is between the upper and lower thresholds, it indicates insufficient statistical evidence to draw conclusions. For the region to be analyzed, the number of heterozygous loci can be continuously accumulated until SPRT classification can be successfully performed.

[0112] The mathematical formulas used to calculate the upper and lower limits of SPRT are: Upper limit threshold = [(ln 8) / N – ln δ] / ln γ; Lower limit threshold = [(ln 1 / 8) / N – ln δ] / ln γ, where δ = (1 – θ1) / (1 – θ2). θ1 is the expected ratio of sequencing tags corresponding to potentially excessive haplotypes when allele imbalance exists in plasma; θ2 is the expected ratio of either haplotype (0.5) when no allele imbalance exists; N is the total number of sequencing tags for Hap I and Hap II. ln The mathematical notation for the natural logarithm, i.e., log e θ1 depends on the relative percentage concentration (F) of tumor DNA in the plasma.

[0113] In the LOH scheme, θ1 = 1 / (2-F). In the single allele amplification scheme, θ1 = (1+zF) / (2+zF), where z represents the additional copy number corresponding to the amplified chromosomal region in the tumor. For example, if a chromosome is replicated, it is an additional copy of that specific chromosome. Then z equals 1.

[0114] Figure 7 This illustration shows an RHDO analysis of plasma DNA from HCC patients targeting a chromosomal segment located on chromosome 1p, according to an embodiment of the invention, wherein the chromosomal segment exhibits monoallelic amplification in tumor tissue. Green triangles represent patient data. The total number of sequencing reads increases with the number of SNPs analyzed. The proportion of total sequencing reads derived from haplotypes corresponding to chromosomal amplification aberrations in the tumor varies with the increase in total analyzed sequencing reads, eventually reaching a value above an upper limit threshold. This indicates a significant haplotype dose imbalance and thus corroborates the presence of this cancer-related chromosomal aberration in the plasma.

[0115] All chromosomal regions in HCC patients were analyzed using SPRT-based RHDO analysis, where these regions showed amplification and deletion in tumor tissue. RHDO analysis was performed separately for 922 chromosomal regions known to have LOH and 105 chromosomal regions known to have amplification, with the following results. For LOH, SPRT was used to classify the 922 chromosomal regions, of which 921 were correctly identified as having haplotype dose imbalance in plasma, providing 99.99% accuracy. For monoallelic amplification, SPRT was used to classify the 105 regions, of which 105 were correctly identified as having haplotype dose imbalance in plasma, providing 100% accuracy.

[0116] C. Analysis of relative haplotype dimensions

[0117] Using the lengths of the corresponding segments from each haplotype can be an alternative method for dose counting of segments corresponding to two haplotypes. For example, for a specific chromosomal region, the size of a DNA segment obtained from one haplotype can be compared to the size of DNA segments from other haplotypes. One can analyze the size distribution of DNA segments corresponding to any allele at the heterozygous locus of the first haplotype in the region and compare it to the size distribution of DNA segments corresponding to any allele at the heterozygous locus of the second haplotype. Statistically significant differences in size distribution can be used to identify aberrations in the same way as dose counting.

[0118] It has been reported that the size distribution of total plasma DNA (i.e., tumor vs. non-tumor DNA) is elevated in cancer patients (Wang BG, et al. Cancer Res. 2003; 63: 3966-8). However, when tumor-derived DNA (not total DNA, i.e., tumor vs. non-tumor DNA) is specifically studied, the length distribution of tumor-derived DNA molecules is observed to be smaller than that of molecules derived from non-tumor cells (Diehl et al. Proc Natl Acad Sci US A. 2005;102:16368-73). Therefore, the size distribution of peripheral blood DNA can be used to determine the presence of cancer-related chromosomal aberrations. The principle of size analysis is shown in... Figure 8 middle.

[0119] Figure 8 The illustration shows a variation in fragment size distribution across two haplotypes of a chromosomal region when a tumor including deletion aberrations is present, according to an embodiment of the invention. Figure 8 As shown, the T allele is deleted in tumor tissue. Consequently, tumor tissue releases only short molecules of the A allele into the plasma. The tumor-derived short DNA molecules cause an overall shortening of the length distribution corresponding to the A allele in the plasma, thus making the length distribution of the A allele in plasma shorter than that of the T allele. As discussed in previous sections, all alleles located on the same haplotype can be analyzed together. In other words, the size distribution of DNA molecules carrying alleles on one haplotype can be compared with the size distribution of DNA molecules carrying alleles on other haplotypes. The haplotype missing in tumor tissue shows a longer size distribution in plasma.

[0120] In addition, size analysis can also be used to detect amplification of chromosomal regions associated with cancer. Figure 9 This illustration shows a variation in fragment size distribution across two haplotypes of a chromosomal region in the presence of a tumor including amplified aberrations, according to an embodiment of the invention. Figure 9 In the example shown, the chromosomal region carrying the T allele is replicated in the tumor. As a result, an increased amount of the shorter DNA molecule carrying the T allele is released into the plasma, thus making the size distribution of the T allele-corresponding fragment appear shorter overall than that of the A allele-corresponding fragment. Similarly, all alleles located on the same haplotype can be analyzed together. In other words, the size distribution of haplotypes amplified in tumor tissue appears shorter than the size distribution of haplotypes not amplified in the tumor.

[0121] Detection of shortened size distribution of peripheral blood DNA

[0122] The size of DNA fragments derived from the two haplotypes (i.e., Hap I and Hap II) can be determined by, but is not limited to, paired-end massively parallel sequencing. After sequencing the ends of the DNA fragment, the sequencing reads (tags) can be aligned to a human reference genome. The size of the sequenced DNA molecule can be derived from the coordinates of the outermost nucleotides at each end. The sequencing tags of the molecule can be used to determine whether the sequenced DNA fragment originated from Hap I or Hap II. For example, one type of sequencing tag may include a heterozygous locus in the chromosomal region to be analyzed.

[0123] Therefore, for each sequenced molecule, we can determine its length and whether it was derived from Hap I or Hap II. Based on the fragment size aligned to its respective haplotype, the computer system can calculate the size distribution (e.g., average fragment size) of Hap I and Hap II. Appropriate statistical analyses can be used to compare the size distributions of DNA fragments derived from Hap I and Hap II to determine if the size distributions are sufficiently different to identify the presence or absence of aberrations. Besides paired-end massively parallel sequencing, other methods can be used to determine the size of DNA fragments, including, but not limited to, sequencing whole DNA fragments, mass spectrometry, and optical methods for observing and comparing the length of observed DNA molecules with standards.

[0124] Next, we introduce two example methods for determining the shortening of peripheral blood DNA associated with gene aberrations in tumors. The aim of these two methods is to provide a quantitative measurement of differences in the size distribution of DNA fragments from two populations. The DNA fragments from the two populations refer to DNA molecules corresponding to Hap I and Hap II.

[0125] Differences in the proportion of short DNA fragments

[0126] In one implementation, the proportion of short DNA fragments is used. A length threshold (w) is set to define short DNA molecules. Different length thresholds can be changed and selected to fit different diagnostic purposes. A computer system can determine the number of molecules that are equal to or shorter than the length threshold. The proportion of short DNA fragments (Q) can then be calculated by dividing the number of short DNA molecules by the total number of DNA fragments. The Q value is affected by the length distribution of the population of DNA molecules. A shorter overall size distribution indicates a higher proportion of short DNA molecules, thus yielding a higher Q value.

[0127] Next, the difference in the corresponding proportions of short DNA fragments between Hap I and Hap II can be used. For Hap I and Hap II, the difference in the proportions of short fragments can reflect the difference in the length distribution of DNA fragments obtained from Hap I and Hap II (ΔQ). ΔQ = Q HapI – Q HapII Q HapI This represents the proportion of short fragments corresponding to the Hap I DNA fragment, while Q... HapII This represents the proportion of short fragments corresponding to Hap II DNA fragments. Q HapI and Q HapII Examples of statistical values ​​for the segment length distribution corresponding to each of the two haplotypes.

[0128] As shown in the previous section, when Hap II is deleted in tumor tissue, the length distribution of the Hap I DNA fragment is shorter than that of the Hap II DNA fragment. As a result, a positive ΔQ value is observed. This positive ΔQ value can be compared to a threshold to determine if the ΔQ is large enough to confirm the presence of deletion. Amplification of Hap I also shows a positive ΔQ value. When Hap II replication is present in tumor tissue, the length distribution of the Hap II DNA fragment is smaller than that of the Hap I DNA fragment. Therefore, the ΔQ value is called negative. In the absence of chromosomal aberrations, the length distributions of the Hap I and Hap II DNA fragments in plasma / serum are similar. Therefore, the ΔQ value is approximately zero.

[0129] The patient's ΔQ can be compared to that of a healthy individual to determine whether the value is normal. Furthermore, the patient's ΔQ value can be compared to values ​​obtained from patients with similar cancers to determine whether the value is abnormal. Such comparisons may involve comparisons with the thresholds described herein. In disease monitoring, the value ΔQ can be continuously monitored over a period of time. Changes in the ΔQ value can indicate an increase or decrease in the relative percentage concentration of tumor DNA in plasma / serum. In selected embodiments of this technique, the relative percentage concentration of tumor DNA can be related to the stage of the tumor, the prognosis of the disease, and its progression. The measurement methods used in such embodiments at different time points will be discussed in more detail below.

[0130] Difference in the proportion of total length contributed by short DNA fragments

[0131] In this embodiment, the proportion of the total length contributed by short DNA fragments is used. The computer system can determine the total length of a set of DNA fragments in a sample (e.g., fragments derived from a specific haplotype of a given region or fragments derived only from a given region). Different cutoff sizes (w) can be selected; below a given cutoff size, DNA fragments are defined as "short fragments." Different cutoff sizes can be varied and selected to fit different diagnostic purposes. The computer system then determines the total length of the short DNA fragments by summing the lengths of randomly selected DNA fragments that are equal to or smaller than the cutoff size. The proportion of the total length contributed by the short DNA fragments is then calculated as follows: F = ∑ w length / ∑ N length, where ∑ w length represents the sum of the lengths of DNA fragments with a length equal to or less than w (bp); while ∑ N "Length" refers to the sum of the lengths of DNA fragments that are equal to or less than a predetermined length N. In one embodiment, N is 600 bases. However, other length constraints, such as 150 bases, 180 bases, 200 bases, 250 bases, 300 bases, 400 bases, 500 bases, and 700 bases, can be used to calculate the "total length".

[0132] Because the Illumina Genome Analyzer system is not very efficient at amplifying and sequencing DNA fragments longer than 600 bases, N can be chosen to be 600 bases. Furthermore, limiting the analysis to DNA fragments shorter than 600 bases avoids measurement bias due to genomic structural variations. In the presence of structural variations, such as rearrangements (Kidd JM et al, Nature 2008; 453:56-64), estimating the length of DNA fragments by aligning the ends of DNA fragments to a reference genome using bioinformatics may overestimate the fragment length. Moreover, of all DNA fragments successfully aligned to the reference genome, more than 99.9% are shorter than 600 bases; therefore, including all fragments equal to and less than 600 bases will provide an unbiased estimate of the length distribution of DNA fragments in the sample.

[0133] Therefore, the difference in the proportion of total length contributed by short DNA fragments between Hap I and Hap II can be used. Variations in the length distribution between Hap I and Hap II DNA fragments can be reflected by their F-values. Here, we will use the F-values... Hap I and F Hap IIThese are defined as the proportions of the short DNA fragments corresponding to Hap I and Hap II, respectively. The difference (ΔF) in the proportions of short DNA fragments between Hap I and Hap II can be calculated as follows: ΔF = F Hap I – F Hap II F Hap I and F Hap II These are the statistical values ​​of the two sets of segment length distributions corresponding to each haplotype.

[0134] Similar to the implementation schemes described above, when comparing Hap I DNA fragments with Hap II DNA fragments, deletion of Hap II in tumor tissue results in an apparent shortening of the Hap I DNA length distribution. This leads to a positive ΔF value. When Hap II is replicated, a negative ΔF value is observed. In the absence of chromosomal aberrations, the ΔF value is close to zero.

[0135] The ΔF value of a patient can be determined to be normal by comparing it to that of a healthy individual. The ΔF value of a patient can be determined to be abnormal by comparing it to that of a patient with a similar cancer. Such comparisons may involve comparisons with the thresholds described herein. In disease monitoring, the ΔQ value can be continuously monitored. Changes in the ΔF value can indicate an increase or decrease in the percentage concentration of tumor DNA in plasma / serum.

[0136] D. General Method

[0137] Figure 10 According to embodiments of the present invention, a flowchart is shown of a method for analyzing the haplotype of a biological sample of an organism to determine whether a chromosomal region exhibits deletion or amplification. The biological sample comprises nucleic acid molecules (also called fragments) derived from normal cells and potentially cancer-related cells. These molecules may be free in the sample. The organism may be any type having more than one copy of chromosomes, i.e., at least a diploid organism, but may include higher polyploid organisms.

[0138] In one embodiment of this and any other method described herein, the biological sample contains cell-free DNA fragments. While plasma DNA analysis is used to illustrate the different methods described in this application, these methods can also be applied to detect tumor-related chromosomal aberrations in samples comprising a mixture of normal and tumor-derived DNA. Other sample types include saliva, tears, pleural fluid, ascites, bile, urine, serum, pancreatic juice, feces, and cervical smear samples.

[0139] In step 1010, at the first chromosomal region, first and second haplotypes are determined for normal cells of the organism. Haplotypes can be determined by any suitable method, such as those mentioned herein. Chromosomal regions can be selected by any method, such as those described herein. The first chromosomal region includes a first plurality of loci (e.g., locus 420 in region 410), and it is heterozygous. Heterozygous loci (hets) can be spaced apart from each other; for example, in the first plurality of loci, one locus may be 500 or 1000 base pairs (or more) apart from another locus. Other heterozygous loci (hets) may be present in the first chromosomal region, but are not necessarily utilized.

[0140] In step 1020, each nucleic acid molecule in the biological sample is characterized by its location and alleles. For example, the location of a nucleic acid molecule within a reference genome of an organism can be identified. This localization can be achieved in various ways, including using molecular sequencing (e.g., via universal sequencing) to obtain one or both (paired-end) sequencing tags of the molecule, and then aligning these sequencing tags to the reference genome. Such alignments can be performed using tools such as the Basic Local Similarity Search (BLAST). This localization can be identified as a specific value within a chromosome arm. Alleles at a heterozygous locus can be used to identify which haplotype a fragment originates from.

[0141] In step 1030, based on the identified location and determined alleles, the first group of nucleic acid molecules is identified as originating from the first haplotype. For example, including... Figure 4 The fragment at locus 421 (which has allele A) shown was identified as derived from Hap I. The first group of nucleic acid molecules can cover the first chromosomal region as long as it includes at least one nucleic acid molecule located at each of the first plurality of loci.

[0142] In step 1040, based on the identified locations and determined alleles, the second group of nucleic acid molecules is identified as originating from the second haplotype. For example, including... Figure 4 The fragment at locus 421 (which has allele T) shown was identified as derived from Hap II. This second group includes at least one nucleic acid molecule located at each of the first plurality of loci.

[0143] In step 1050, the computer system calculates a first value for the first group of nucleic acid molecules. This first value defines the properties of the first group of nucleic acid molecules. Examples of the first value include the tag count corresponding to the number of nucleic acid molecules in the first group and the length distribution corresponding to the number of nucleic acid molecules in the first group.

[0144] In step 1060, the computer system calculates a second value for the second group of nucleic acid molecules. This second value defines the properties of the second group of nucleic acid molecules.

[0145] In step 1070, the first value is compared with the second value to determine whether the first chromosomal region exhibits a classification of deletion or amplification. The presence of a classification of deletion or amplification can provide information about whether an organism has cancer-associated cells. Examples of comparisons include considering the difference or ratio between two values, and comparing the result to one or more thresholds, as described herein. For example, in SPRT analysis, the ratio can be compared to a threshold. Instance classifications can include positive (i.e., detected amplification or deletion), negative, undetermined, and the degree of variation corresponding to positive and negative (e.g., using integers from 1 to 10, or real numbers from 0 to 1). Amplification can include simple replication. Such methods can detect the presence of cancer-associated nucleic acids, including tumor DNA and DNA derived from pre-tumor lesions (i.e., precancerous lesions).

[0146] E. Depth

[0147] The depth of analysis refers to the number of molecules required for analysis to provide classification or other assays that meet specific accuracy requirements. In one embodiment, the depth can be calculated based on known distortions, and then measurements and analyses can be performed under conditions that satisfy that depth. In another embodiment, analysis can be performed continuously until successful classification is achieved, and the depth corresponding to successful classification can be used to determine the level of cancer (e.g., cancer stage or tumor size). Some examples of depth-related calculations are provided below.

[0148] As described in this article, bias can refer to any difference or ratio. For example, bias can lie between a first and a second value derived from a threshold or tumor concentration, or between a first parameter and a second parameter. If the bias is doubled, the number of fragments that need to be measured is reduced to 1 / 4. More generally, if the bias increases by a factor of N, the number of fragments that need to be measured is 1 / N of the original number. 2 Therefore, if the deviation is reduced to 1 / N, the number of segments to be tested will increase to N times the original number. 2 N can be a real number or an integer.

[0149] Consider a scenario where tumor DNA constitutes 10% of the sample (e.g., plasma), and assuming sequencing yields 10 million fragments with statistically significant differences observed. Then, for example, an enrichment process is performed so that the sample contains 20% tumor DNA; the required number of fragments would be 2,500,000. In this way, depth can be correlated with the percentage concentration of tumor DNA in the sample.

[0150] Furthermore, the amount of amplification also affects depth. For a region whose copy number is amplified to twice its original normal copy number (e.g., 4, compared to the normal 2), assume that X number of fragments are needed for analysis. If the copy number of that region is increased to 4 times its original copy number, then that region requires X / 4 number of fragments.

[0151] F. Threshold

[0152] As described above, the deviation or bias of parameters from normal values ​​(e.g., differences or ratios between haplotypes) can be used to provide a diagnosis. For example, the bias can be the difference between the average size of fragments obtained from one haplotype of a region and the average size of fragments obtained from other haplotypes. If the bias is greater than a certain amount (e.g., a threshold determined from normal samples and / or regions), deletions or amplifications are identified. However, the degree above the threshold can be useful, so multiple thresholds can be used, each corresponding to a different level of cancer. For example, a higher bias from normal values ​​can indicate the stage of cancer (e.g., stage 4 has a higher degree of imbalance than stage 3). Furthermore, a higher bias could also be due to a larger tumor releasing many fragments, and / or the analyzed region being amplified multiple times.

[0153] In addition to providing different levels of cancer, varying thresholds allow for the efficient detection of regions with aberrations or specific areas. For example, a high threshold can be set to primarily look for amplifications of 3x or higher, which yields a greater imbalance than a single haplotype deletion. Furthermore, deletions in regions with two copies can be detected. Conversely, lower thresholds can be used to identify potentially aberrant regions, which can then be further analyzed to determine the presence and location of the aberration. For instance, a binary search (or a higher-level search, such as an octree search) can be used, employing a higher threshold at lower levels within the search hierarchy.

[0154] Figure 11 This illustration shows a subregion 1130 with deletion aberrations within region 1110 of a cancer cell, and measurements taken in plasma to determine the deleted region. Chromosomal regions 1110 can be selected by any method mentioned herein, for example, by dividing the genome into segments of equal size. Furthermore, regarding each locus of locus 1120, Figure 11 The number of allele counts is also shown. In addition, the cumulative totals are recorded for region 1140 (normal region) and region 1130 (deleted region).

[0155] If region 1110 is selected for analysis, the cumulative count for Hap I is 258, while the cumulative count for Hap II is 240, resulting in a difference of 18 across 11 loci. This difference represents a much smaller percentage of the total counts than if only the corresponding subregion 1130 were analyzed. This is reasonable, as approximately half of region 1110 is normal, whereas in cancer cells, the entire subregion 1130 is deleted. Therefore, aberrations in region 1110 may be missed, depending on the threshold used.

[0156] To allow for the detection of deletions in sub-regions, multiple implementations can use lower thresholds for relatively large regions (in this example, assuming region 1110 is relatively larger than the region to be deleted). Lower thresholds will identify more regions, potentially including some false positives, but they will reduce false negatives. False positives can now be removed through further analysis, which can also precisely locate distortions.

[0157] Once a region is identified for further analysis, it can be divided into different sub-regions for further analysis. Figure 11 In this study, 11 loci can be divided in half (e.g., using a binary tree), resulting in a subregion 1140 of 6 loci and a subregion 1130 of 5 loci. These regions can be analyzed using the same or a more stringent threshold. In this example, subregion 1140 is then identified as normal, while subregion 1130 is identified as containing deletions or amplifications. In this way, larger regions can be rejected as not having aberrations, and time can be spent further analyzing suspicious regions (regions above the lower threshold) to identify subregions showing aberrations with high reliability (e.g., using a higher threshold). Although this paper uses RHDO, the sizing techniques are equally applicable.

[0158] The size of the region used for the first-level search (and the size of the sub-regions at lower levels in the tree) can be selected based on the size of the region to be detected and the size of the distortion. Studies have found that cancer exhibits 10 distortion regions with a length of 10 MB. Furthermore, some patients have distortion regions as large as 100 MB. Later-stage cancers can have even larger distortion regions.

[0159] G. Improvement of distortion location within the region

[0160] In the previous section, we discussed tree-search-based partitioning of regions into subregions. Here, we discuss other methods for analyzing subregions and how to accurately locate distortions within those regions.

[0161] Figure 12This illustration shows how RHDO analysis can be used to map the location of aberrations according to an embodiment of the invention. Chromosomal regions are shown horizontally, with haplotypes in non-cancerous cells labeled as Hap I and Hap II. Regions lacking Hap II in cancerous cells are labeled as LOH.

[0162] As shown, the RHDO analysis begins to the left and proceeds to the right of the hypothetical chromosome region 1202. Each arrow represents a chromosomal segment for RHDO classification. Each chromosomal segment can be considered its own region, specifically a subregion of a larger heterozygous region. The size of the RHDO classification segment depends on the number of loci (and their locations) before classification is determined. The number of loci included in each RHDO segment depends on the number of molecules used for the segment analysis, the required reliability (e.g., the odds ratio in SPRT), and the relative percentage concentration of tumor-derived DNA in the sample. Figure 4 and Figure 5 In the example shown, classification is performed when the number of molecules is sufficient to determine a statistically significant difference between the two haplotypes.

[0163] Each solid horizontal arrow represents an RHDO classification segment, indicating the absence of haplotype dose imbalance in the DNA sample. In regions of the tumor without LOH, six RHDO classifications are formed, each representing the absence of haplotype dose imbalance. The next RHDO classification segment 1210 crosses the junction 1205 between regions with and without LOH. Figure 12 The lower section shows the SPRT curves for RHDO segment 1210. The black vertical arrows indicate the junctions between regions with and without LOH. As more data is accumulated from regions with LOH, the RHDO classification of this chromosomal segment indicates a haplotype dose imbalance.

[0164] Each white horizontal arrow represents an RHDO classification segment, indicating the presence of a haplotype dose imbalance. Furthermore, the next four RHDOs on the right indicate the presence of a haplotype dose imbalance in the DNA sample. It can be inferred that the binding point between regions with and without LOH is located in the first segment showing a change in RHDO classification, i.e., from the presence of a haplotype dose imbalance to the absence of a haplotype dose imbalance, and vice versa.

[0165] Figure 13 The classification of RHDOs starting from another direction is shown according to an embodiment of the invention. Figure 13The diagram illustrates RHDO classification starting from two directions. RHDO analysis starting from the left can infer that the binding point between regions with and without LOH is located within the first RHDO segment 1310, which shows the presence of a haplotype dose imbalance. RHDO analysis starting from the right can infer that the binding point is located within the first RHDO segment 1320, which shows the absence of a haplotype dose imbalance. Combining the information from RHDO analyses performed from both directions, the location of the binding point 1330 between regions with and without LOH can be inferred.

[0166] IV. Detection of nonspecific haplotypes of aberrations

[0167] The RHDO method relies on the use of heterozygous loci. Now, diploid organisms have chromosomes with some differences, resulting in two haplotypes, but the number of heterozygous loci varies. Some individuals may have relatively few heterozygous loci. Furthermore, the implementation described in this section can also be used for homozygous loci, where two haplotypes are compared between two regions rather than the same region. Therefore, although there may be some disadvantages due to comparing two different chromosomal regions, more data points can be obtained.

[0168] In the relative chromosomal region dosing method, the number of fragments obtained from a chromosomal region (e.g., determined by counting sequencing tags aligned to that region) is compared to a desired value (which may be obtained from a reference chromosomal region or from the same region in another known healthy sample). In this approach, fragments can be counted for chromosomal regions regardless of the haplotype from which the sequencing tags are derived. Therefore, sequencing tags that do not contain heterozygous sites can still be used. For comparison purposes, the implementation scheme can normalize the tag counts before comparison. Each region is defined by at least two loci (the two loci are separated from each other), and fragments at these loci can be used to obtain a total value for the region.

[0169] For a specific region, a normalized value for the sequencing reads (tags) can be calculated by dividing the number of sequencing reads aligned to that region by the number of sequencing reads aligned to the entire genome. This normalized tag count allows for comparison of results obtained from one sample with results from another. For example, the normalized value can be the expected proportion (e.g., percentage or fraction) of sequencing reads obtained from a specific region, as described above. However, many other normalizations are possible, as will be apparent to those skilled in the art. For instance, one can normalize by dividing the count of a region by the count of a reference region (in the above case, the reference region is the entire genome). This normalized tag count can then be compared to a threshold, which can be determined by one or more reference samples that do not have cancer.

[0170] Next, the normalized label count of the test individual is compared with the normalized label counts of one or more reference subjects (e.g., those without cancer). In one implementation, the comparison is made by calculating a z-score for the test individual targeting a specific chromosomal region. The z-score is calculated using the following equation: z-score = (Normalized label count in the case – Mean) / SD, where “mean” is the average of the normalized label counts mapped to a specific chromosomal region for the reference sample; and SD is the standard deviation of the normalized label counts mapped to the specific region for the reference sample. Therefore, the z-score is the number of standard deviations relative to the standard deviation, i.e., how many standard deviations away the normalized label count of the test individual is from the mean of the normalized label counts of one or more reference subjects for a given chromosomal region.

[0171] In organisms with cancer, amplified chromosomal regions in tumor tissue are present in excess in plasma DNA. This results in a positive z-score. Conversely, deleted chromosomal regions in tumor tissue are present in insufficient amounts in plasma DNA. This results in a negative z-score. The magnitude of the z-score depends on a variety of factors.

[0172] One factor is the relative percentage concentration of tumor-derived DNA in a biological sample (e.g., plasma). A higher relative percentage concentration of tumor-derived DNA in a sample (e.g., plasma) results in a greater difference between the normalized tag count of the tested individual and the count of a reference individual. Therefore, a larger z-score is obtained.

[0173] Another factor is the variability of normalized label counts in one or more reference individuals. In biological samples (e.g., plasma) of the individuals being tested, smaller variability (i.e., smaller standard deviation) of normalized label counts in the reference group will result in a larger z-score when the degree of excess in chromosomal regions is the same. Similarly, in biological samples (e.g., plasma) of the tested cases, smaller standard deviation of normalized label counts in the reference group will result in a larger negative z-score in absolute value when the degree of under-exposition in chromosomal regions is the same.

[0174] Another factor is the degree of chromosomal aberration in tumor tissue. For a specific chromosomal region, the degree of chromosomal aberration refers to the change in copy number (increase or loss). The greater the copy number change in tumor tissue, the greater the degree of over- or under-presentation of the specific chromosomal region in plasma DNA. For example, the loss of two copies of a chromosome results in a greater degree of under-presentation of the chromosomal region in plasma than the loss of one of the two copies of a chromosome, thus yielding a larger negative z-score. Typically, multiple chromosomal aberrations are present in cancer. In various cancers, chromosomal aberrations can further vary in their nature (i.e., amplification or deletion), their degree (single or multiple copy increases or losses), and their length (size of the aberration).

[0175] The accuracy of measuring normalized tag counts is affected by the number of molecules analyzed. We anticipate that 15,000, 60,000, and 240,000 molecules need to be analyzed at percentage concentrations of approximately 12.5%, 6.3%, and 3.2% to detect chromosomal aberrations (additions or losses) with a one-copy change. Details of tag counts for detecting cancer targeting different chromosomal regions are described in U.S. Patent Publication No. 2009 / 0029377 (Lo et al.) entitled “Diagnosing Fetal Chromosomal Aneuploidy Using Massively Parallel Genomic Sequencing,” the entire contents of which are incorporated herein by reference for all purposes.

[0176] Furthermore, several embodiments may use length analysis instead of tag counting. Additionally, length analysis may be used without normalized tag counts. As mentioned herein and described in U.S. Patent Application No. 12 / 940,992, length analysis can use multiple parameters. For example, Q or F values ​​obtained above can be used. Since such length values ​​are not proportional to the number of readings, these values ​​do not need to be normalized by counts obtained from other regions. Techniques of haplotype-specific methods can also be applied to non-specific methods. For example, techniques involving region depth and modifications can be used. In some embodiments, when comparing two regions, a correction based on the GC content of a specific region may be considered. Since the RHDO method uses the same regions, such a correction is unnecessary.

[0177] V. Multiple regions

[0178] While some cancers can typically present with specific chromosomal regions, these cancers do not always present only in the same regions. For example, other chromosomal regions may show aberrations, and the location of these other regions may be unknown. Furthermore, when screening for cancer to identify early-stage cancers, it is desirable to differentiate between multiple cancers that may present with aberrations anywhere throughout the genome. To address these situations, multiple implementations can be used to systematically analyze multiple regions to determine which regions show aberrations. For example, the number of aberrations and their location (e.g., whether they are contiguous) can be used to confirm aberrations, the stage of cancer, provide a cancer diagnosis (e.g., whether the number exceeds a certain threshold), and provide a prognosis based on the number and location of multiple regions presenting aberrations.

[0179] Therefore, multiple implementations can identify whether an organism has cancer based on the number of regions displaying aberrations. Thus, one can examine multiple regions (e.g., 3000) to identify the number of regions displaying aberrations. These regions can cover the entire genome or only a portion of the genome, such as non-repeating regions.

[0180] Figure 14 The flowchart illustrates a method 1400 for analyzing biological samples of an organism using multiple chromosomal regions, according to an embodiment of the present invention. The biological sample includes nucleic acid molecules (also referred to as fragments).

[0181] In step 1410, multiple non-overlapping chromosomal regions of the organism are identified. Each chromosomal region comprises multiple loci. As mentioned above, the size of the region can be 1 Mb, or some other equivalent size. The entire genome can comprise approximately 3000 regions, each with a predetermined size and location. Furthermore, as mentioned above, such predetermined regions can be varied to accommodate a specific length of a particular chromosome or a specific number of regions to be used, as well as any other criteria mentioned herein. If regions have different lengths, such lengths can be used to normalize the results, for example, as described herein.

[0182] In step 1420, for each of the plurality of nucleic acid molecules, its position within the reference genome of the organism is identified. Positions can be determined in any manner described herein, such as by obtaining sequencing tags from sequencing fragments and aligning the sequencing tags with the reference genome. Furthermore, specific haplotypes of the molecules can be detected from and can be obtained using haplotype-specific methods.

[0183] Steps 1430-1450 are performed for each chromosomal region. In step 1430, the chromosomal region corresponding to each group of nucleic acid molecules is identified based on the identified location. Each group of nucleic acid molecules includes at least one nucleic acid molecule located at each of multiple loci within the chromosomal region. In one embodiment, the group can be a fragment corresponding to a specific haplotype of the chromosomal region, for example, as described in the RHDO method above. In another embodiment, the group can be any fragment corresponding to the chromosomal region, for example, as described in part IV.

[0184] In step 1440, the computer system calculates values ​​for each group of nucleic acid molecules. Each value defines the properties of each group of nucleic acid molecules. Each value can be any value mentioned herein. For example, the value can be the number of fragments in the group or a statistical value of the length distribution of the fragments in the group. The value can also be a normalized value, such as the label count of a region divided by the total number of label counts in the sample, or divided by the number of label counts in a reference region. Furthermore, the value can also be a difference or ratio (e.g., in RHDO), thereby providing properties corresponding to different regions.

[0185] In step 1450, each value is compared to a reference value to determine whether the first chromosomal region exhibits a classification of deletion or amplification. This reference value can be any threshold or reference value described herein. For example, the reference value can be a threshold determined for a normal sample. For RHDO, each value can be a difference or ratio of tag counts for two haplotypes, and the reference value can be a threshold used to determine the presence of a statistically significant difference. For example, the reference value can be a tag count or size for another haplotype or region, and the comparison can include, but is not limited to, a difference or ratio (or a function of such values), and then it is determined whether the difference or ratio is greater than a threshold.

[0186] The reference value can be changed based on the results of other regions. For example, if adjacent regions also show bias (albeit small compared to a threshold, such as a z-score of 3), a lower threshold can be used. For instance, if three consecutive regions are all above a first threshold, the probability of cancer is higher. Therefore, this first threshold can be lower than another threshold necessary for identifying cancer by non-contiguous regions. Having three (or more) regions with small biases can reduce the likelihood of the influence of random fluctuations sufficiently to maintain high sensitivity and specificity.

[0187] In step 1460, the number of chromosomal regions classified as exhibiting deletions or amplifications is determined. There may be some limitations on the chromosomal regions counted. For example, only regions contiguous with at least one other region may be counted (or contiguous regions may be required to reach a certain size, such as four or more regions). In embodiments where the regions are not equal, the number may consider their respective lengths (e.g., the number may be the total length of the aberrant regions).

[0188] In step 1470, the quantity is compared to a threshold to determine the classification of the sample. For example, classification could be whether the organism has cancer, the stage of cancer, and the prognosis of the cancer. In one embodiment, all aberrant regions are counted using a single threshold, regardless of where the regions appear. In another embodiment, the threshold can be varied based on the location and size of the counted regions. For example, the quantity of a region on a specific chromosome or chromosome arm can be compared to a threshold for that specific chromosome (or arm). Multiple thresholds can be used. For example, the quantity of aberrant regions on a specific chromosome (or arm) must be greater than a first threshold, and the total number of aberrant regions in the genome must be greater than a second threshold.

[0189] The threshold for the quantity of a region can also depend on how strong the imbalance is in the regions being counted. For example, the quantity of a region used to determine the threshold for cancer classification can depend on the specificity and sensitivity of measuring distortion in each region (distortion threshold). For example, if the distortion threshold is low (e.g., a z-score of 2), a high quantity threshold (e.g., 150) can be chosen. However, if the distortion threshold is high (e.g., a z-score of 3), the quantity threshold can be lower (e.g., 50). Furthermore, the quantity of a region showing distortion can also be a weighted value; for example, regions showing high imbalance (i.e., those with more classifications than just distortion-positive and negative) can be given a higher weight than regions showing less imbalance.

[0190] Therefore, the quantity of chromosomal regions (which may include number and / or size) can be used to reflect the severity of the disease, where the chromosomal region refers to a corresponding normalized tag count (or values ​​corresponding to other group attributes) showing a significant over- or under-presentation. The quantity of chromosomal regions with aberrations indicated by normalized tag counts can be determined by two factors: the number (or size) of chromosomal aberrations in tumor tissue and the percentage concentration of tumor-derived DNA in a biological sample (e.g., plasma). More advanced cancers tend to show more (and larger) chromosomal aberrations. Therefore, larger cancer-associated chromosomal aberrations are potentially detectable. In patients with more advanced cancers, a higher tumor burden leads to a higher percentage concentration of tumor-derived DNA in the plasma. As a result, tumor-associated chromosomal aberrations are more easily detected in plasma samples.

[0191] In cancer screening or detection, the quantity of chromosomal regions can be used to determine the likelihood of a tested subject having cancer. The chromosomal region refers to the region whose corresponding normalized label count (or other value) shows a significant over- or under-presentation. Using a cutoff of ±2 (i.e., z-score >2 or <-2), it is expected that approximately 5% of the tested regions will show a significant deviation from the mean of control subjects due to random fluctuations in z-score. When the entire genome is divided into 1 Mb segments, there are approximately 3000 chromosomal regions. Therefore, approximately 150 chromosomal regions are expected to have z-scores >2 or <-2 (due to random fluctuations).

[0192] Therefore, a threshold of 150 can be used to determine the presence of cancer for the number of segments with z-scores >2 or <-2. For the number of segments with aberration z-scores (e.g., 100, 125, 175, 200, 250, and 300), other cutoff values ​​can be chosen to fit the diagnostic purpose. Lower cutoff values ​​(e.g., 100) result in higher sensitivity but lower specificity, while higher cutoff values ​​result in higher specificity but lower sensitivity. The number of false positives can be reduced by increasing the z-score cutoff value. For example, if the cutoff value is increased to 3, only 0.3% of segments are false positives. In this case, more than three segments with aberration z-scores can be used to indicate the presence of cancer. Furthermore, other thresholds, such as 1, 2, 4, 5, 10, 20, and 30, can be chosen to fit different diagnostic purposes. However, as the number of aberration fragments required for diagnosis increases, the sensitivity of detecting cancer-related chromosomal aberrations decreases.

[0193] A feasible method to improve sensitivity without sacrificing specificity would consider the results of adjacent chromosomal segments. In one implementation, the cutoff value for the z-score is kept between >2 and <-2. However, a chromosomal region is classified as potentially aberrant only when two consecutive segments show the same type of aberration, for example, both segments have z-scores >2. If the bias of the normalized label count is random error, the probability of having two consecutive segments (in the same direction) as a false positive is 0.125% (5% x 5% / 2). On the other hand, if the chromosomal aberration covers two consecutive segments, a lower cutoff value will make the detection of over- or under-presentation of segments in plasma samples more sensitive. Since the bias of the normalized label count (or other value) from the mean of control subjects is not due to random error, the requirement for consecutive classification does not have a significantly adverse effect on sensitivity. In other implementations, if a higher cutoff value is used, the z-scores of adjacent segments can be summed. For example, the z-scores of three consecutive segments can be summed, and a cutoff value of 5 can be used. This concept can be extended to more than three consecutive segments.

[0194] Furthermore, the combination of quantity and aberration threshold can also depend on the purpose of the analysis and any prior knowledge (or lack thereof) about the organism. For example, in cancer screening of a normal healthy population, a high specificity corresponding to the quantity of regions (i.e., a high threshold used in terms of the quantity of regions) and an aberration threshold are typically used to identify whether a region has a chromosomal aberration. However, in patients at higher risk (e.g., those with a mass or family history, who smoke, have HPV, hepatitis virus, or other viruses), the threshold can be lowered, resulting in higher sensitivity (lower false negatives).

[0195] In one implementation, if tumor-derived DNA with a resolution of 1 Mb and a detection limit of 6.3% is used to detect chromosomal aberrations, then the number of molecules in each 1 Mb segment must be 60,000. For the entire genome, this would require approximately 180,000,000 (60,000 readings / Mb x 3,000 Mb) comparable readings.

[0196] Figure 15 The following diagram illustrates, according to an embodiment of the invention, the depth required to elucidate different numbers of chromosomal segments and the relative percentage concentrations of tumor-derived fragments in Table 1500. Column 1510 provides the concentration of DNA fragments obtained from the tumor cells of the sample. Higher concentrations facilitate the detection of aberrations, thus requiring fewer molecules for analysis. Column 1520 provides an estimated number of molecules required for each chromosomal segment, which can be calculated using the methods described above in the section on depth.

[0197] Smaller chromosomal segments offer higher resolution, suitable for detecting smaller chromosomal aberrations. However, this requires an increased number of molecules for analysis. Larger chromosomal segments, at the cost of reduced resolution, reduce the number of molecules needed for analysis. Therefore, only larger aberrations can be detected. In one implementation, the larger the region used, the more subdivided the aberrant chromosomal segment will be, and these sub-regions will be analyzed, resulting in higher resolution (e.g., as described above). Column 1530 provides the size of each chromosomal segment. Smaller values ​​correspond to a larger number of regions. Column 1540 shows the number of molecules required for the entire genome. Therefore, if one has the evaluation parameters (or minimum detection concentration) described above, the number of molecules required for analysis can be determined.

[0198] VI. Progress over a certain period of time

[0199] As tumors progress, the amount of tumor DNA fragments in the plasma increases due to the release of more DNA fragments (e.g., due to tumor growth, increased gangrene, or higher vascularity). The increased amount of DNA fragments corresponding to tumor tissue entering the plasma increases the degree of imbalance in the plasma (e.g., in RHDO, the difference in tag counts between two haplotypes will increase). Furthermore, due to the increased number of tumor DNA fragments, the number of regions containing aberrations can be more easily detected. For example, the amount of tumor DNA in a region might be so small that chromosomal aberrations cannot be detected. This is because in cases where the tumor is small and only a small amount of cancer DNA fragments are released, there are not enough fragments for analysis, so a statistically significant difference cannot be established, making the aberration undetectable. Even in small tumors, more fragments may be available for analysis, but this may require a large sample size (e.g., a large amount of plasma).

[0200] The progression of cancer can be tracked using the amount of aberration in one or more regions (e.g., by imbalance or desired depth), or the amount (number and / or size) of chromosomal regions exhibiting aberration. In one instance, if the amount of aberration in one region (or regions) increases more rapidly than in other regions, that region can be used as a preferred molecular marker for cancer detection. This increase may be a result of tumor enlargement and / or the release of more DNA fragments due to repeated amplification of the region. Furthermore, changes in postoperative aberration values ​​(e.g., the amount of aberration or the number of regions exhibiting aberration, or a combination thereof) can be monitored to confirm whether the tumor has been completely removed.

[0201] In various embodiments of the technology, determining the relative percentage concentration of tumor DNA can be used for cancer staging, prognosis, or monitoring cancer progression. The measured progression can provide information about the current stage of the cancer and how quickly it is growing or spreading. Cancer “staging” is related to all or some of the following factors: tumor size, histological appearance, presence / absence of lymph node metastasis, and presence / absence of distant metastasis. Cancer “prognosis” involves assessing the probability of disease progression and / or the probability of survival from cancer. Furthermore, it involves the assessment of time during which the patient may not have clinical progression or survival time. Cancer “monitoring” involves seeing if the cancer has progressed (e.g., increased size, lymph node metastasis, or spread to distant organs, i.e., metastasis). Additionally, monitoring may involve checking whether the tumor is under treatment control. For example, if treatment is effective, one may see a reduction in tumor size, regression of metastasis or lymph node metastasis, and an improvement in the patient’s overall health (e.g., weight gain).

[0202] A. Determination of the relative percentage concentration of cancer DNA

[0203] One method for tracking an increase in the amount of aberration in one or more regions is to determine the relative percentage concentration of cancer DNA in that region. The tumor can then be tracked over a period of time using changes in the relative percentage concentration of cancer DNA. This tracking can be used for diagnosis; for example, the first measurement can provide a background level (which may correspond to the general aberration level in a person), while subsequent measurements, if showing changes, indicate tumor growth (and therefore cancer). Furthermore, changes in the relative percentage concentration of cancer DNA can be used to evaluate the prognostic effect of treatment. In other embodiments of the technique, an increase in the relative percentage concentration of tumor DNA in plasma indicates a poorer prognosis or an increased tumor burden in the patient.

[0204] The relative percentage concentration of cancer DNA can be determined in several ways. For example, in tag counting, the difference between one haplotype and another (or one region compared to another). Another approach is to look at the depth prior to statistically significant differences (i.e., the number of fragments analyzed). An example of the former is that haplotype dose differences can be used to determine the relative percentage concentration of tumor-derived DNA in biological samples (e.g., plasma) by analyzing chromosomal regions with loss of heterozygosity.

[0205] Previous studies have shown a positive correlation between the amount of tumor-derived DNA and tumor burden in cancer patients (Lo et al. Cancer Res. 1999;59:5452-5. and Chan et al. Clin Chem. 2005;51:2192-5). Therefore, continuous monitoring of the relative percentage concentration of tumor-derived DNA in biological samples (e.g., plasma samples) via RHDO analysis can be used to monitor the progression of a patient's disease. For example, monitoring the relative percentage concentration of tumor-derived DNA in samples collected sequentially after treatment (e.g., plasma) can be used to determine the success of treatment.

[0206] Figure 16 The principle of measuring the relative percentage concentration of tumor-derived DNA in plasma by RHDO analysis according to an embodiment of the present invention is illustrated. The imbalance between two haplotypes is determined, and the degree of imbalance can be used to determine the relative percentage concentration of tumor DNA in the sample.

[0207] Hap I and Hap II represent two haplotypes in non-tumor tissues. Hap II is partially deleted in subregion 1610 in tumor tissues. Therefore, the Hap II-related fragments detected in plasma corresponding to the deleted region 1610 are contributed by non-tumor tissues. On the other hand, region 1610 in Hap I is present in both tumor and non-tumor tissues. Therefore, the difference between Hap I and Hap II readings represents the amount of tumor-derived DNA in plasma.

[0208] The relative percentage concentration of tumor-derived DNA can be calculated using the following formula for chromosomal regions affected by LOH, based on the number of sequencing reads (tags) obtained from deleted and non-deleted chromosomes: F = (N HapI –N HapII ) / N HapI x 100%, where N HapI For heterozygous SNPs located in LOH-affected chromosomal regions, this refers to the number of sequencing reads corresponding to the alleles on Hap I; while N HapII For heterozygous SNPs located in LOH-affected chromosomal region 1610, the number of sequencing reads corresponding to the alleles in Hap II.

[0209] The above formula is equivalent to defining p as the cumulative tag count of heterozygous loci located in chromosomal regions excluding deletions (Hap I), and defining q as the cumulative tag count located in chromosomal regions including deletion 1610 (Hap II), and the relative percentage concentration (F) of tumor DNA in the sample is calculated as F = 1 - q / p. Figure 11 In the example shown, the relative percentage concentration of tumor DNA was 14% (1 – 104 / 121).

[0210] Plasma samples A, representing a specific percentage concentration of tumor DNA, were collected from HCC patients before and after tumor resection. Prior to tumor resection, for a given chromosomal region, N... HapI The value is 30443, while for the second haplotype of the chromosomal region, N HapII The value was 16221, resulting in an F value of 46.7%. After tumor resection, N... HapI It is 31534, and N HapII The value was 31098, and the F-value was 1.4%. This monitoring indicates that the tumor resection was successful.

[0211] The degree of change in the size distribution of free DNA can also be used to determine the relative percentage concentration of tumor DNA. In one embodiment, the exact size distributions of DNA derived from both tumor and non-tumor tissues in plasma can be determined, and the measured size distribution falling between two known distributions can provide the percentage concentration of tumor DNA (e.g., using a linear model between two statistical values ​​of the size distributions of tumor and non-tumor tissues). Furthermore, continuous monitoring of size changes can also be used. In one aspect, the change in size distribution is determined to be proportional to the percentage concentration of tumor DNA in plasma.

[0212] Furthermore, differences between different regions can be used in a similar manner, namely, the non-specific haplotype detection methods described above. In tag counting methods, multiple parameters are used to monitor disease progression. For example, regarding regions showing chromosomal aberrations, the magnitude of the z-score can be used to reflect the percentage concentration of tumor-derived DNA in a biological sample (e.g., plasma). The degree of over- or under-presentation of a particular region is proportional to the percentage concentration of tumor-derived DNA in the sample or the degree or amount of copy number variation in tumor tissue. The magnitude of the z-score is a parameter that measures the degree of over- or under-presentation of a particular chromosomal region in a sample compared to a control subject. Therefore, the magnitude of the z-score can reflect the relative percentage concentration of tumor DNA in the sample, and thus reflect the patient's tumor burden.

[0213] B. Number of tracking areas

[0214] As mentioned above, the number of regions exhibiting chromosomal aberrations can be used to screen for cancer and also for monitoring and prognosis. For example, this monitoring can be used to determine the current stage of the cancer if it recurs or if treatment is effective. As a tumor progresses, its genomic composition is further degraded. To identify this continuous degradation, methods that track the number of regions (e.g., predetermined regions of 1 Mb) can be used to identify the progression of the tumor. In more advanced cancers, the tumor will have even more regions exhibiting aberrations.

[0215] C. Method

[0216] Figure 17According to embodiments of the present invention, a flowchart is shown of a method for determining the process of chromosomal aberrations in an organism using a biological sample comprising nucleic acid molecules. In one embodiment, at least some of the nucleic acid molecules are free-floating. For example, chromosomal aberrations may originate from malignant tumors or pre-existing lesions. Furthermore, an increase in aberrations may be due to an increasing number of cells with chromosomal aberrations in the organism over a certain period of time, or due to an increase in the number of aberrations per cell in a certain proportion of the organism's cells. As an example of reduction, treatments (e.g., surgery or chemotherapy) can remove or reduce cancer-related cells.

[0217] In step 1710, one or more non-overlapping chromosomal regions of the organism are identified. Each chromosomal region comprises multiple gene loci. Regions can be identified by any suitable method, such as those described herein.

[0218] Steps 1720-1750 are performed at each of the multiple time points. Each time point corresponds to a different time when the sample is obtained from the organism. The current sample is the sample to be analyzed during a given time period. For example, a sample can be taken once a month for 6 months and analyzed as soon as the sample is obtained. Alternatively, multiple measurements can be taken and analyzed over multiple time periods.

[0219] In step 1720, the current biological sample of the organism is analyzed to identify the location of the nucleic acid molecule within the organism's reference genome. The location can be determined in any manner mentioned herein, such as by sequencing fragments to obtain sequencing tags and aligning the sequencing tags to the reference genome. Furthermore, specific haplotypes of the molecule can be determined using haplotype-specific methods.

[0220] Steps 1730-1750 are performed for each of one or more chromosomal regions. When multiple regions are used, the implementation scheme obtained from Part V can be used. In step 1730, the chromosomal region corresponding to each group of nucleic acid molecules can be identified based on the identified location. Each group of nucleic acids includes at least one nucleic acid molecule located at each of a plurality of loci within the said chromosomal region. In one implementation, the group can be a fragment aligned to a specific haplotype of the chromosomal region, such as that described in the RHDO method above. In another implementation, the group can be any fragment aligned to the chromosomal region, as described in Part IV.

[0221] In step 1740, the computer system calculates values ​​for each group of nucleic acid molecules. Each value defines a property of each group of nucleic acid molecules. Each value can be any value mentioned herein. For example, the value can be the number of fragments in the group, or a statistical value of the size distribution of the fragments in the group. Furthermore, each value can be a normalized value, such as, for a sample, the tag count of a region divided by the total number of tag counts, or divided by the number of tag counts of a reference region. Additionally, each value can be a difference or ratio to another value (e.g., in RHDO), thereby providing a property of difference for the region.

[0222] In step 1750, each value is compared to a reference value to determine whether the first chromosomal region exhibits a deletion or amplification. This reference value can be any threshold or reference value described herein. For example, the reference value can be a threshold determined for a normal sample. For RHDO, each value can be the difference or ratio of tag counts for two haplotypes, and the reference value can be a threshold used to determine the presence of a statistically significant difference. As another example, the reference value can be the tag count or size of another haplotype or region, and the comparison can include, but is not limited to, differences or ratios (or functions of such values), and then it is determined whether the difference or ratio is greater than a threshold. The reference value can be determined according to any suitable method and standard, such as those described herein.

[0223] In step 1760, the classification of individual chromosomal regions at multiple time points is used to determine the nature of chromosomal aberrations in the organism. This process can be used to determine whether an organism has cancer, the stage of cancer, and the prognosis of cancer. Each of these determinations may involve the classification of cancer, as described herein.

[0224] This type of cancer classification can be implemented in several ways. For example, the amount of aberrant regions can be counted and compared to a threshold. In terms of regions, the classification can be numerical (e.g., tumor concentration, with reference values ​​for different haplotypes or different regions), and changes in concentration can be measured. These changes can be compared to a threshold, and if a significant increase is detected, this indicates the presence of a tumor.

[0225] VII. Examples

[0226] A. Using SPRT-based RHDO

[0227] In this section, we present an example of analysis using the relative haplotype dose (RHDO) of SPRT in a patient with hepatocellular carcinoma (HCC). In this patient's tumor tissue, deletion of one of the two chromosome 4 segments was observed. This results in a loss of heterozygosity for the SNP on chromosome 4. Regarding the patient's haplotype, the genomic DNA of the patient, his wife, and son was analyzed, and the genotypes of the three individuals were determined. The patient's structural haplotype was derived from their genotypes. Massive parallel sequencing was performed, and sequencing reads with SNP alleles corresponding to the two haplotypes on chromosome 4 were identified and counted.

[0228] The equations and principles of RHDO and SPRT have been described above. In one embodiment, RHDO analysis can be computer-programmed to detect, for example, a 10% difference in haplotype dosage in a DNA sample, where the dosage difference corresponds to 10% of the tumor DNA concentration when one of two haplotypes is amplified or deleted. In other embodiments, the sensitivity of RHDO analysis can be set to detect 2%, 5%, 15%, 20%, 25%, 30%, 40%, and 50% of tumor-derived DNA in a DNA sample. The sensitivity of RHDO analysis can be tuned using parameters that calculate the upper and lower thresholds of the SPRT classification region. For example, using the ratio ratio (the ratio of the tag count of one haplotype to the tag count of other haplotypes), adjustable parameters can be the desired detection limit level (e.g., what percentage of tumor concentration should be detectable, which will affect the number of molecules analyzed) and the threshold used for classification.

[0229] In this RHDO analysis, the null hypothesis was that the two haplotypes on chromosome 4 were present at the same dose. The alternative hypothesis was that the doses of the two haplotypes differed by more than 10% in a biological sample (e.g., plasma). For both hypotheses, the number of sequencing reads with SNP alleles (corresponding to the two haplotypes) was statistically compared in the form of data accumulated from different SNPs. SPRT classification was performed when the accumulated data were sufficient to determine whether the doses of the two haplotypes were present at the same dose or statistically differed by at least 10%. A typical SPRT classification region on the q-arm of chromosome 4 is shown below. Figure 18A The 10% threshold is used here for illustrative purposes only. Other levels of difference can also be detected (e.g., 0.1%, 1%, 2%, 5%, 15%, or 20%). Generally, the lower the level of difference one wants to detect, the more DNA molecules need to be analyzed. Conversely, the higher the level of difference one wants to detect, the fewer DNA molecules need to be analyzed to obtain a statistically significant difference. For this analysis, the odds ratio is used for SPRT, but other parameters such as z-scores or p-values ​​can be used.

[0230] In plasma samples from HCC patients obtained at diagnosis, 76 and 148 successful RHDO classifications were found for the p and q arms of chromosome 4, respectively. All RHDO classifications indicated haplotype dose imbalances in the plasma samples obtained at diagnosis. For comparison, plasma samples from patients obtained after surgical resection of the tumor were also analyzed. Figure 18B As shown. For post-treatment samples, there were 4 and 9 successful RHDO classifications for the p and q arms of chromosome 4, respectively. All four RHDO classifications showed no observable haplotype dose imbalance corresponding to a tumor DNA concentration greater than 10% in the plasma samples. Of the nine RHDO classifications for chromosome 4q, seven showed a lack of haplotype dose imbalance, while two showed an imbalance. The number of RHDO blocks showing dose imbalance corresponding to a tumor DNA concentration greater than 10% decreased significantly after tumor resection, indicating that the size of the chromosomal region showing dose imbalance was significantly smaller in post-treatment samples than in pre-treatment samples. These results suggest that the relative percentage concentration of tumor DNA in plasma decreases after surgical tumor resection.

[0231] When compared with non-haplotype-specific methods, RHDO analysis provides a more accurate estimate of the relative percentage concentration of tumor DNA and is particularly well-suited for monitoring the progression of aberrations. Therefore, one can expect disease progression to be characterized by an increase in the percentage concentration of tumor DNA in the plasma; conversely, if the patient's disease is stable or their tumor has regressed or decreased in size, the percentage concentration of tumor DNA in the plasma will decrease.

[0232] B. Targeted Analysis

[0233] In some alternative implementations, target sequence capture technology can be combined with universal sequencing of DNA fragments. This method is also referred to herein as targeted analysis. One implementation of this method involves preferentially selecting fragments using a liquid-phase capture system (e.g., the AgilentSureSelect system, the Illumina TruSeq Custom Enrichment Kit (illumina.com / applications / sequencing / targeted_resequencing.ilmn), or via the MyGenostics GenCapCustom Enrichment system (mygenostics.com / )) or a microarray-based capture system (e.g., the Roche NimbleGene system). While some other regions can be captured, certain regions are preferred. Such methods allow for the analysis of these regions at higher depths (e.g., more fragments can be sequenced or analyzed using digital PCR) and / or at lower costs. Greater depth can increase sensitivity within the region. Other enrichment methods can be implemented based on fragment size and methylation patterns.

[0234] Therefore, an alternative approach to whole-genome DNA sample analysis is to target specific regions of interest to detect common chromosomal aberrations. Because the analysis focuses on regions with potential chromosomal aberrations, regions exhibiting specific tumor-related changes, or regions of specific clinical importance, targeted analysis could potentially improve the cost of this method. Examples of such changes as described above include those occurring early in tumorigenesis of specific cancer types (e.g., in HCC, the presence of amplifications at 1q and 8q and deletions at 8q are early chromosomal changes - van Malenstein et al. Eur J Cancer 2011;47:1789-97); or changes associated with prognosis (e.g., increases at 6q and 17q and deletions at 6p and 9p observed during tumor progression, and in colorectal cancer patients, the presence of LOH at 18q, 8p, and 17p is associated with poorer survival - Westra et al. Clin Colorectal Cancer 2004;4:252-9); or changes that can predict a response to treatment (e.g., an increase in the presence at 7p predicts a response to tyrosine kinase inhibitors in patients with epidermal growth factor receptor mutations - Yuan et al. J Clin Oncol 2011;29:3435-42). Other examples of altered genomic regions in cancer can be found in numerous online databases (e.g., the Cancer Genome Anatomy Project database (cgap.nci.nih.gov / Chromosomes / RecurrentAberrataions) and the Atlas of Genetics and Cytogenetics in Oncology and Haematology (atlasgeneticsoncology.org / / Tumors / Tumorliste.html)). Conversely, in non-targeted whole-genome approaches, regions where chromosomal aberrations are unlikely to occur are analyzed to the same degree (depth) as regions with potential aberrations.

[0235] We used a targeted analysis strategy to analyze plasma samples from three HCC patients and four healthy controls. Target enrichment was achieved using the SureSelect capture system from Agilent (Gnirke et al. Nat. Biotechnol 2009.27:182-9). The SureSelect system was chosen as an example of a feasible target enrichment technique. Other liquid-phase (Illumina TruSeq Custom Enrichment system) or solid-phase (e.g., Roche-Nimblegen system) target capture systems, as well as amplicon-based target enrichment systems (e.g., QuantaLife system and RainDance system), can also be used. Capture probes were designed to locate chromosomal regions showing aberrations, both common and rare in HCC. After target capture, DNA samples were sequenced one lane per flow cell in an Illumina GAIIx analyzer. Regions with very low probability of amplification and deletion were used as references for comparison with regions where amplification and deletion were relatively frequent.

[0236] exist Figure 19 The figure shows common chromosomal aberrations found in HCC (adapted from Wong et al. (Am J Pathol 1999;154:37-43)). Lines on the right side of the chromosome diagram represent chromosome increases, while lines on the left represent chromosome deletions in individual patient samples. Thick lines indicate high levels of increases. Rectangles indicate the locations of target-capturing probes.

[0237] Target tag counting analysis

[0238] For the detection of chromosomal aberrations, we first calculate normalized tag counts for potential aberration regions and reference regions. Then, as described previously by Chen et al. (PLoS One 2011;6:e21791), the normalized tag counts are corrected for the GC content of the region. In this example, the p-arm of chromosome 8 is selected as the potential aberration region, and the q-arm of chromosome 9 is selected as the reference region. Using an Affymetrix SNP 6.0 microarray, tumor tissues from three HCC patients were analyzed for chromosomal aberrations. The changes in chromosomal doses of 8p and 9q in the tumor tissues of the three patients are shown below. Patient HCC013 showed a decrease in 8p and no change in 9q. Patient HCC027 showed an increase in 8p and no change in 9q. Patient HCC023 showed a decrease in 8p and no change in 9q.

[0239] Next, using targeted analysis, the ratio of normalized label counts between chr 8p and 9q was calculated for 3 HCC patients and 4 healthy controls. Figure 20A The results for the normalized tag count ratios in HCC and healthy patients are shown. For the cases HCC013 and HCC023, a decrease in the normalized tag count ratio between 8p and 9q was observed. This is consistent with the discovery of chromosome 8p deletion in the tumor tissue. For the case HCC027, an increased ratio was observed, and this increased ratio is consistent with the increase in chromosome 8p in the tumor tissue of this condition. The dashed lines represent deviations of plus or minus two standard deviations from the mean of the four normal control subjects.

[0240] Target size analysis

[0241] In the previous section, we described the principle of identifying cancer-related alterations by determining the size distribution of plasma DNA fragments in cancer patients. Furthermore, target enrichment methods can be used to detect size changes. For three HCC examples (HCC 013, HCC027, and HCC023), the size of each sequenced DNA fragment was determined after aligning the sequencing reads to a human reference genome. The size of the sequenced DNA fragment was derived from the coordinates of the two outermost nucleotides at the ends. In other embodiments, the entire length of the DNA fragment is sequenced, and the fragment size can then be determined directly from the sequencing length. The size distribution of DNA fragments aligned to chromosome 8p was compared with that aligned to chromosome 9q. For detecting differences in the size distribution of DNA between the two populations, the proportion of DNA fragments shorter than 150 bp was first determined for each population in the current examples. In other implementations, other cutoff values ​​can be used, such as 80 bp, 110 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 160 bp, and 170 bp. ΔQ = Q 8p – Q 9q Q 8p This refers to the proportion of DNA fragments aligned to chromosome 8p that are shorter than 150 bp; while Q 9q The proportion of DNA fragments that are aligned to chromosome 9q and shorter than 150 bp.

[0242] Because shorter DNA fragment sizes result in a higher proportion of fragments shorter than the cutoff value (in this example, a proportion shorter than 150 bp), a higher (larger positive) ΔF value indicates that DNA fragments aligned to chromosome 8p have a shorter distribution compared to those aligned to chromosome 9q. Conversely, a smaller (or larger negative) result indicates that DNA fragments aligned to chromosome 8p have a longer distribution compared to those aligned to chromosome 9q.

[0243] Figure 20B The results of size analysis following target enrichment and massively parallel sequencing are shown for 3 HCC patients and 4 healthy controls. In the 4 healthy controls, a positive ΔQ value indicates that DNA fragments aligned to chromosome 8p have a slightly shorter size distribution than those aligned to chromosome 9q. The dashed lines represent the intervals corresponding to ΔQ values ​​within two standard deviations from the mean for the 4 controls. The ΔQ values ​​for cases HCC 013 and HCC023 are shifted downwards by more than two standard deviations relative to the control mean. Both cases have a deletion of chromosome 8p in the tumor tissue. For this chromosomal region, the deletion of 8p in the tumor leads to a reduced contribution of tumor-derived DNA to the plasma. Since tumor-derived DNA is shorter in peripheral blood than DNA derived from non-tumor tissues, this results in a significantly longer size distribution of plasma DNA fragments aligned to chromosome 8p. This is consistent with the lower (larger negative) ΔQ values ​​in these two examples. In contrast, in the case of HCC027, amplification of 8p results in a significantly shorter distribution of DNA fragments aligned to that region. Therefore, it is assumed that the higher proportion of plasma DNA fragments aligned to 8p are relatively short. This is consistent with observations that the ΔQ value for HCC027 has a larger absolute positive value than that of healthy controls.

[0244] C. Multiple regions used to detect tumor-derived chromosomal aberrations

[0245] Chromosomal aberrations, including deletions and amplifications of certain chromosomal regions, are commonly detected in tumor tissue. Characteristic patterns of chromosomal aberrations are observed in different types of cancer. Here, we use several examples to illustrate different methods for detecting these cancer-related chromosomal aberrations in the plasma of cancer patients. Furthermore, our method is also suitable for cancer screening, monitoring disease progression, and response to treatment. Samples from one HCC patient and two nasopharyngeal carcinoma (NPC) patients were analyzed. For the HCC patient, venous blood samples were collected before and after surgical resection of the tumor. For the two NPC patients, venous blood samples were collected at diagnosis. In addition, plasma samples from one chronic hepatitis B carrier and one subject whose plasma contained detectable Epstein-Barr virus DNA were analyzed. Neither of these subjects had any cancer.

[0246] Microarray analysis was used to detect tumor-derived chromosomal aberrations. Specifically, the Affymetrix SNP6.0 microarray system was used to analyze DNA extracted from blood cells and tumor samples from HCC patients. The Affymetrix Genotyping Console v4.0 was used to determine the genotypes of blood cells and tumor tissues. The Birdseedv2 algorithm was used to determine chromosomal aberrations, including additions and deletions, based on the signal intensity of different alleles of SNPs on the microarray and copy number alteration (CNV) probes.

[0247] Count-based analysis

[0248] To perform sequencing tag counting analysis in plasma, 10 mL of venous blood was collected from each subject. Plasma was collected after centrifugation of each blood sample. DNA was extracted from 4–6 mL of plasma using the QIAmp blood mini Kit (Qiagen). Plasma DNA libraries were constructed as previously described (Lo YMD. Sci Transl Med 2010, 2:61ra91) and then massively parallel sequencing was performed on the Illumina Genome Analyzer platform. End-paired sequencing was performed for plasma DNA analysis. Each molecule was sequenced (50 bp) at each of its two ends, for a total of 100 bp per molecule. Using the SOAP2 program (soap.genomics.org.cn / ) (Li R et al. Bioinformatics 2009, 25:1966-7), the two ends of each sequence were aligned to the human genome (Hg18 NCBI.36 downloaded from UCSC genome.ucsc.edu).

[0249] The genome was then divided into multiple 1-megabase (1Mb) chromosomal segments, and the number of sequencing reads aligned to each 1Mb segment was determined. Next, based on the GC content of each chromosomal segment, a locally weighted regression scatter smoothing (LOESS) algorithm was used to correct the tag counts for each segment (Chen E et al. PLoS One 2011, 6:e21791). The purpose of this correction was to minimize sequencing-related quantitative bias, which is caused by differences in GC content between different genomic segments. The 1Mb segmentation mentioned above is for illustrative purposes. Other segment sizes can also be used, such as 2 Mb, 10 Mb, 25 Mb, 50 Mb, etc. Furthermore, segment sizes can be selected based on the genomic characteristics of a specific tumor in a specific patient and specific types of tumors in general. Furthermore, for example, with single molecular sequencing technologies, if the sequencing method can show low GC bias, such as the Helicos system (www.helicosbio.com) or the Pacific Biosciences Single Molecular Real-Time system (www.pacificbiosciences.com), the GC correction step can be omitted.

[0250] In a previous study, we sequenced 57 plasma samples from subjects who did not have cancer. The results of these plasma sequencing were used to determine a reference range for tag counts in each 1Mb segment. For each 1Mb segment, the mean and standard deviation of the tag counts for the 57 individuals were determined. The results for the study subjects were then expressed as z-scores, which were calculated using the following equation: z-score = (number of sequencing tags in the case – mean) / SD, where the mean is the average number of sequencing tags aligned to a specific 1Mb segment for the reference samples; and SD is the standard deviation of the sequencing tags aligned to a specific 1Mb segment for the reference samples.

[0251] Figure 21-24 The results of sequencing tag count analysis for four study participants are shown. 1Mb regions are indicated at the edges of the figure. Chromosome numbering and chromosome pattern diagrams (outermost circle) are adjusted clockwise from pter-qter (centromeres are shown in yellow). Figure 21In the diagram, the inner circle 2101 shows the regions of aberrations (deletions or amplifications) determined by tumor analysis. The inner circle 2101 is shown in five levels, ranging from -2 (the innermost line) to +2 (the outermost line). A value of -2 indicates the loss of two chromosome copies for the corresponding region. For a given chromosomal region, a value of -1 indicates the absence of one of the two chromosome copies. A value of 0 indicates no chromosome addition or deletion. A value of +1 indicates the addition of one chromosome copy, while +2 indicates the addition of two chromosome copies.

[0252] The middle circle 2102 shows the results of the plasma analysis. As can be seen, the distortion results in the plasma before treatment in the middle circle 2102 are consistent with the distortion pattern in the inner circle. The middle circle 2102 is a line with more levels, but the process is the same. The outer circle 2103 shows the data points obtained from the analysis of the plasma after treatment, and these data points are grayed out (indicating no excessive or insufficient presentation - no distortion).

[0253] Chromosomal regions with excessive sequencing tag presentation in plasma (z-score > 3) are represented by green dots 2110. Regions with insufficient sequencing tag presentation in plasma (z-score < -3) are represented by red dots 2120. Regions with no significant chromosomal aberrations detected in plasma (z-score between -3 and 3) are represented by gray dots. The total number of excessive / insufficient presentation pairs is normalized. In cases where PCR amplification was involved prior to sequencing, normalization can account for GC bias correction.

[0254] Figure 21 A Circos diagram of an HCC patient according to an embodiment of the present invention is shown, depicting data obtained from the counting of sequencing tags of plasma DNA. The circles, from the inside out, represent: chromosomal aberrations detected in tumor tissue by microarray analysis (red and green represent deletions and amplifications, respectively); z-score analysis of plasma samples obtained before surgical resection of the tumor; and z-score analysis of plasma samples obtained one month after resection. Prior to tumor resection, the chromosomal aberrations detected in the plasma precisely matched those identified in the tumor tissue by microarray analysis. After tumor resection, most cancer-related chromosomal aberrations disappeared from the plasma. These data reflect the value of such methods for detecting disease progression and treatment efficacy.

[0255] Figure 22 This illustration shows a counting analysis of sequencing tags performed on plasma samples from chronic HBV carriers who do not have HCC, according to an embodiment of the invention. Compared with HCC patients ( Figure 21 Conversely, cancer-related chromosomal aberrations were not detected in the plasma of this HBV carrier. These data reflect the value of this method for cancer screening, diagnosis, and surveillance.

[0256] Figure 23 This paper illustrates a counting analysis of sequencing tags performed on plasma samples from cancer patients with stage III NPC according to an embodiment of the present invention. Chromosomal aberrations were detected in the plasma samples obtained prior to treatment. Specifically, significant aberrations were identified in chromosomes 1, 3, 7, 9, and 14.

[0257] Figure 24 This illustration shows a sequencing tag counting analysis of plasma samples from a cancer patient with stage IV NPC according to an embodiment of the present invention. Chromosomal aberrations were detected in the plasma samples obtained before treatment. When compared with patients with stage III cancer (… Figure 23 Compared to patients with [specific disease type], more chromosomal aberrations were detected. Furthermore, the sequencing tag counts deviated more from the control mean, i.e., the z-scores deviated more from 0 (positive or negative). The increased number of chromosomal aberrations and the greater degree of deviation in sequencing tag counts compared to controls reflect a more severe degree of genomic alteration in advanced disease, thus highlighting the value of such methods for cancer staging, prognosis, and surveillance.

[0258] Size-based analysis

[0259] Previous studies have shown that the size distribution of DNA derived from tumor tissue is shorter than that of DNA derived from non-tumor tissue (Diehl F et al. Proc Natl Acad Sci USA 2005, 102(45):16368-73). In previous studies, we summarized a method for detecting plasma haplotype imbalances through plasma DNA size analysis. Here, we use sequencing data from HCC patients to further illustrate this method.

[0260] For illustrative purposes, we identified two regions for size analysis. In one region (chromosome 1 (chr1); coordinates: 159,935,347 to 167,219,158), a duplicate of one of the two homologous chromosomes was detected in tumor tissue. In another region (chromosome 10 (chr10); coordinates: 100,137,050 to 101,907,356), a deletion (LOH) of one of the two homologous chromosomes was detected in tumor tissue. In addition to determining which haplotype the sequencing fragment originated from, the size of the sequencing fragment was determined bioinformatically using the coordinates of the outermost nucleotide of the sequencing fragment in the reference genome. Next, the size distribution of the fragments corresponding to each of the two haplotypes was determined.

[0261] For the LOH region of Chr10, a haplotype (i.e., the deleted haplotype) was detected in tumor tissue. Therefore, all plasma DNA fragments matched to this deleted haplotype were derived from non-cancer tissues. On the other hand, fragments matched to the undeleted haplotype (non-deleted haplotype) in tumor tissue could be derived from either tumor or non-tumor tissues. Since the size distribution of tumor-derived DNA is shorter, we predict that the size distribution of fragments corresponding to the non-deleted haplotype will be shorter than that corresponding to the deleted haplotype. The difference between the two size distributions can be determined by plotting the cumulative frequency of fragments against the size of DNA fragments. The DNA population with a shorter size distribution will have a relatively larger number of short DNA fragments; therefore, the cumulative frequency of the DNA population with a shorter size distribution will increase more rapidly within the intervals corresponding to shorter fragments in the size range.

[0262] Figure 25 This diagram illustrates the relationship between the cumulative frequency and size of plasma DNA in regions exhibiting LOH (Left-of-H) in tumor tissue, according to an embodiment of the invention. The X-axis represents the size of fragments in base pairs. The Y-axis represents the percentage of fragments smaller than the values ​​on the X-axis. Compared to sequences obtained from deleted haplotypes, sequences obtained from non-deleted haplotypes show a faster increase and a higher cumulative frequency corresponding to sizes smaller than 170 bp. This indicates that short DNA fragments obtained from non-deleted haplotypes are more abundant. This is consistent with the predictions above, given the size distribution of tumor-derived short DNA obtained from non-deleted haplotypes.

[0263] In one implementation, the difference in size distribution can be measured by the difference in the cumulative frequencies of the two DNA molecule populations. We define ΔQ as the difference in the cumulative frequencies of the two populations. ΔQ = Q 非删除的 – Q 删除的 Q 非删除的 Q represents the cumulative frequency of sequencing DNA fragments obtained from the haplotype corresponding to the non-deletion; while Q... 删除的 This represents the cumulative frequency of the sequenced DNA fragments obtained by deleting the corresponding haplotype.

[0264] Figure 26 The relationship between ΔQ and the size of sequenced plasma DNA in the LOH region is shown. According to an embodiment of the invention, ΔQ reaches 0.2 at a size of 130 bp. This indicates that using 130 bp as the cutoff value for defining short DNA is optimal for use in the equation described above. Using this cutoff value, the short DNA molecular weight is as much as 20% higher in the population derived from the deleted haplotype compared to the population derived from the non-deleted haplotype. This percentage difference (or a similar derived value) is then compared to a threshold derived from individuals without cancer.

[0265] In regions with chromosomal amplification, within tumor tissue, one haplotype is replicated (the amplified haplotype). Because the additional amount of tumor-derived short DNA molecules resulting from this amplified haplotype are released into the plasma, the size distribution of fragments from the amplified haplotype is shorter than that from the non-amplified haplotype. Similar to the LOH protocol, differences in size distribution can be determined by plotting the cumulative frequency of fragments based on their size. A population of DNAs with shorter size distributions contains more short DNA, and therefore their cumulative frequency increases more rapidly within the intervals corresponding to the shorter fragments in the size range.

[0266] Figure 27 The following diagram illustrates the relationship between the cumulative frequency of plasma DNA and the size of DNA fragments corresponding to regions of chromosomal replication in tumor tissue, according to an embodiment of the invention. Sequences obtained from amplified haplotypes show a faster increase in frequency and a higher cumulative frequency corresponding to sizes smaller than 170 bp compared to sequences obtained from non-amplified haplotypes. This indicates that a greater quantity of short DNA fragments is obtained from amplified haplotypes. Since a large amount of tumor-derived short DNA is derived from amplified haplotypes, the above is consistent with expectations presented below.

[0267] Similar to the LOH scheme, differences in size distribution can be quantitatively analyzed by the difference in the cumulative frequencies of DNA molecules between two populations. We define ΔQ as the difference in cumulative frequencies between the two populations. ΔQ = Q 扩增的 – Q 非扩增的 Q 扩增的 Q represents the cumulative frequency of sequencing DNA fragments obtained from amplified haplotypes; while Q... 非扩增的 This indicates the cumulative frequency of sequencing DNA fragments obtained from non-amplified haplotypes.

[0268] Figure 28 The following diagram illustrates the relationship between ΔQ and the size of sequenced plasma DNA for the amplified region, according to an embodiment of the invention. According to an embodiment of the invention, ΔQ reaches 0.08 when the size is 126 bp. This indicates that using 126 bp as the cutoff value for defining short DNA results in a significantly higher molecular weight (up to 8%) in the population obtained from amplified haplotypes compared to the population obtained from non-amplified haplotypes.

[0269] D. Other technologies

[0270] In other implementations, sequence-specific techniques can be used. For example, oligonucleotides can be engineered to hybridize with fragments in specific regions. The oligonucleotides are then counted in a manner similar to sequencing tag counting. This method can be used to detect cancers exhibiting specific aberrations.

[0271] VIII. Computer Systems

[0272] Any computer system mentioned in this article may use any suitable number of subsystems. Examples of such subsystems are shown below. Figure 9 The computer system comprises a single computer instrument (900). In some embodiments, the computer system includes a single computer instrument, wherein a subsystem may be a component of the computer instrument. In other embodiments, the computer system may include multiple computer instruments (each a subsystem) and internal components.

[0273] Figure 29 The subsystems shown are interconnected via system bus 2975. Other subsystems coupled to display adapter 2982 are shown, such as printer 2974, keyboard 2978, hard disk 2979, monitor 2976, etc. External and input / output (I / O) devices (coupled to I / O controller 2971) can be connected to the computer system via any means known in the art (serial port 2977). For example, serial port 2977 or external interface 2981 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 2900 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 2975 allows central processing unit 2973 to communicate with the various subsystems and allows central processing unit 2973 to control system memory 2972 ​​or hard disk 2979 to execute instructions and control information exchange between subsystems. System memory 2972 ​​and / or hard disk 2979 can be computer-readable media. Any value mentioned in this article can be output from one component to another and can be output to the user.

[0274] A computer system may include multiple identical components or subsystems connected together, for example, via an external interface 2981 or an internal interface. In some embodiments, the computer system, subsystem, or instrument may be connected to a network. In this case, one computer may be considered a client and another computer a server, where both may be components of the same computer system. Both the client and the server may include multiple systems, subsystems, or components.

[0275] It should be understood that any embodiment of the present invention can be implemented using hardware and / or computer software in a modular or integrated manner as control logic. Based on the disclosed invention and guidance provided herein, those skilled in the art will understand and appreciate other ways and / or methods of implementing embodiments of the present invention using hardware and combinations of hardware and software.

[0276] Any software component or function described in this application can be implemented as software code (based on conventional or object-oriented techniques) executed by a processor using any suitable computer language (such as Java, C++, or Perl). The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media (such as hard disk drives or floppy disks), or optical media (such as optical discs (CDs) or DVDs (Digital Universal Optical Discs)), flash memory, etc. The computer-readable medium can be any combination of such storage or transmission devices.

[0277] Furthermore, such programs can be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks (compliant with various protocols, including the Internet). Thus, digital signals encoded by such programs can be used to create computer-readable media according to embodiments of the invention. Computer-readable media encoded using such programs can be embedded in compatible devices or provided separately by other devices (e.g., downloaded via the Internet). Any such computer-readable media can reside on or within a single computer program product (e.g., a hard disk drive, CD, or an entire computer system) and can reside on or within different computer program products within a system or network. The computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.

[0278] Any method described herein can be implemented, in whole or in part, using a computer system (including a processor) configured to implement the steps described herein. Therefore, multiple embodiments are possible with respect to a computer system potentially using different components that implement individual steps or groups of steps to implement the steps of any method described herein. Although the methods described herein are presented as numbered steps, the steps of these methods can be implemented simultaneously or in different orders. Furthermore, a portion of these steps can be used in conjunction with a portion of other steps of other methods. Moreover, all or part of the steps can be optional. Furthermore, any step of any method can be implemented using modules, circuits, or other means for implementing these steps.

[0279] Specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the invention. However, other embodiments of the invention may relate to specific embodiments concerning individual aspects or specific combinations of these individual aspects.

[0280] For the purpose of illustrating and describing the invention, exemplary embodiments of the invention have been presented in the foregoing description. It is not intended to be exhaustive or to limit the invention entirely to the forms described in the examples, and many modifications and variations are also possible based on the foregoing description. Several embodiments were chosen and described in order to best illustrate the principles of the invention and its practical application, thereby enabling those skilled in the art to best utilize the invention in various embodiments and various modifications (suitable for the particular purpose contemplated).

[0281] Unless otherwise specified, the description of “a” (“a” or “an”) or “the” (“the”) means “one or more”. For all purposes, all patents, patent applications, published inventions and descriptions mentioned above are incorporated herein by reference in their entirety. None of these documents are identified as prior art.

Claims

1. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to perform a method of determining a cancer level of a subject from a biological sample of the subject, the biological sample comprising a plurality of nucleic acid molecules, the method comprising: identifying a plurality of non-overlapping chromosomal regions in a reference genome of an organism, each chromosomal region comprising a plurality of loci; identifying, for each nucleic acid molecule of the plurality of nucleic acid molecules in the biological sample of the subject, a location of the nucleic acid molecule in the reference genome; for each of the plurality of non-overlapping chromosomal regions: identifying, based on the identified locations, groups of nucleic acid molecules from the chromosomal region, the groups comprising at least one nucleic acid molecule at each locus of the plurality of loci; calculating a value for the groups of nucleic acid molecules, the value defining a property of the groups of nucleic acid molecules; and comparing the value to a first threshold reference value and a second threshold reference value to classify the chromosomal region; wherein the first threshold reference value deviates from a normal value by a first amount; wherein the second threshold reference value deviates from the normal value by a second amount, the second amount being greater than the first amount; wherein the first threshold reference value and the second threshold reference value each deviate from the normal value in the same direction; and wherein if the value exceeds the first threshold reference value but does not exceed the second threshold reference value, the chromosomal region is assigned a first classification; and if the value exceeds the first threshold reference value and the second threshold reference value, the chromosomal region is assigned a second classification; and determining the cancer level of the subject based on the classification for each of the plurality of non-overlapping chromosomal regions.

2. The computer readable medium of claim 1, wherein the first classification corresponds to a first stage of cancer and the second classification corresponds to a second stage of cancer, wherein the second stage of cancer is a higher stage than the first stage of cancer.

3. The computer readable medium of claim 1, wherein the first classification corresponds to a first tumor size and the second classification corresponds to a second tumor size, wherein the second tumor size is greater than the first tumor size.

4. The computer readable medium of claim 1, wherein the first threshold reference value is derived from one or more healthy organisms or from one or more regions that do not have deletions or amplifications.

5. The computer readable medium of claim 1, wherein each of the plurality of non- overlapping chromosomal regions is 500 Kb - 2 Mb in length.

6. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to perform a method of identifying one or more chromosomal aberrations of an organism from a biological sample of the subject, the method comprising: (a) identifying a chromosomal region corresponding to a reference genome of the organism, the chromosomal region comprising a plurality of loci; (b) identifying, for each nucleic acid molecule of the plurality of nucleic acid molecules in the biological sample of the subject, a location of the nucleic acid molecule in the reference genome; (c) identifying, based on the identified locations, a first set of nucleic acid molecules as being from the chromosomal region; (d) calculating a first value for the first set of nucleic acid molecules, the first value defining a property of the first set of nucleic acid molecules; (e) comparing the first value to a first threshold value, the first threshold value being offset from a first normal value by a first amount; (f) as a result of the first value exceeding the first threshold value, identifying a second set of nucleic acid molecules as being from a sub-region within the chromosomal region, wherein the second set of nucleic acid molecules is a subset of the first set of nucleic acid molecules; (g) calculating a second value for the second set of nucleic acid molecules, the respective value defining a property of the second set of nucleic acid molecules; (h) comparing the second value to a second threshold value, the second threshold value being offset from a second normal value by a second amount, the second amount being of a higher order than the first amount; (i) identifying, based on the comparison of the second value to the second threshold value, whether the sub-region of the chromosomal region exhibits a aberration.

7. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to perform a method of analyzing a biological sample of an organism, the biological sample comprising nucleic acid molecules derived from normal cells and potentially cancer-related cells, wherein at least some of the nucleic acid molecules are free in the biological sample, the method comprising: identifying one or more chromosomal regions of the organism, each first chromosomal region comprising a first plurality of loci and being associated with a copy number aberration of cancer; enriching nucleic acid molecules from the one or more first chromosomal regions of the biological sample, thereby providing an enriched biological sample; for each of a plurality of nucleic acid molecules in the enriched biological sample of the organism: identifying a location of the nucleic acid molecule in a reference genome of the organism; and for each first chromosomal region of the one or more chromosomal regions: identifying, based on the identified locations, a first set of nucleic acid molecules as being from the first chromosomal region, the first set comprising at least one nucleic acid molecule located at each locus of the first plurality of loci of the first chromosomal region; calculating a first value for the first set of nucleic acid molecules, the first value defining a property of the first set of nucleic acid molecules; wherein the first value determines a parameter; and comparing the parameter to a reference value, thereby determining a classification of whether the first chromosomal region exhibits a deletion or an amplification in any cancer-related cells.

8. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to perform a method of analyzing a biological sample of an organism, wherein the biological sample is from a biological fluid comprising nucleic acid molecules derived from normal cells and potentially tumor cells, and wherein at least some of the nucleic acid molecules are free in the biological sample, the method comprising: causing sequencing or sequence-specific probes to run on a plurality of nucleic acid molecules in the biological sample of the organism to obtain sequencing reads; receiving the sequencing reads at a computer system; identifying a chromosomal region comprising a plurality of loci; ​ ​ for each of the plurality of nucleic acid molecules that are free in the biological sample of the organism: identifying, using the sequencing reads, a location of the nucleic acid molecule in a reference genome of the organism, identifying, based on the identified location, a first set of nucleic acid molecules as being from a chromosomal region of the reference genome, the first set comprising at least one nucleic acid molecule located at each of a plurality of loci of the chromosomal region; using the computer system, calculating a respective value for the first set of nucleic acid molecules, the respective value defining an amount or size of the first set of nucleic acid molecules; comparing the respective value to a reference value to determine a classification of whether the chromosomal region exhibits a deletion or an amplification; determining a difference between the respective value and the reference value; and based on the difference, determining a value for the concentration of the portion of DNA of tumor origin.

9. A computer readable medium storing instructions that, when executed by a computer system, cause the computer system to implement a method of analyzing a biological sample of an organism, wherein the biological sample is from a biological fluid comprising nucleic acid molecules originating from both normal cells and potentially tumor cells, and wherein at least some of the nucleic acid molecules are free in the biological sample, the method comprising: (a) running sequencing or sequence-specific probes on a plurality of nucleic acid molecules in the biological sample of the organism to obtain sequencing reads; (b) receiving the sequencing reads at a computer system; (c) identifying a chromosomal region comprising a plurality of loci, the chromosomal region identified as comprising a deletion or an amplification; (d) for each of the plurality of nucleic acid molecules that are free in the biological sample of the organism: identifying, using the sequencing reads, a location of the nucleic acid molecule in a reference genome of the organism, (e) identifying, based on the identified location, a first set of nucleic acid molecules as being from a first sub-region of the chromosomal region, the first set comprising at least one nucleic acid molecule located at each of a first subset of the plurality of loci; (f) increasing the plurality of loci in the first subset until a comparison of a respective value for the first set of nucleic acid molecules to a reference value indicates a first classification about whether the chromosomal region exhibits a deletion or an amplification with a specified statistical accuracy; (g) continuing to analyze other sub-regions of the chromosomal region comprising other subsets of loci by repeating step (f) to obtain other classifications for the other sub-regions; and (h) identifying a junction point of the amplification or deletion based on changes in the other classifications from one sub-region to another sub-region.

10. A system comprising: the computer readable medium of any one of claims 1-9; and one or more processors for executing the instructions stored on the computer readable medium. ​

Citation Information

Patent Citations

  • Eccentrically-adjustable cam for scales

    US1636873A

  • Diagnosing fetal chromosomal aneuploidy using massively parallel genomic sequencing

    US20090029377A1

  • Fetal Genomic Analysis From A Maternal Biological Sample

    US20110105353A1

  • Size-based genomic analysis

    US20110276277A1