Enhancement of cancer screening using cell-free viral nucleic acids

By analyzing methylation patterns of cell-free viral DNA, the method improves tumor detection sensitivity and specificity, effectively distinguishing between healthy and tumor-bearing individuals and reducing false results.

JP2025186258APending Publication Date: 2025-12-23THE CHINESE UNIVERSITY OF HONG KONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025137553
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-07-26
Filing Date
2025-08-21
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Current methods for detecting tumors early in their development lack sensitivity and specificity, often resulting in false-positive or false-negative results due to the low prevalence of tumors and small amounts of tumor material in samples, particularly when viral DNA is present without cancer.

Method used

Analyze cell-free DNA molecules in biological samples by determining the methylation status of viral DNA, comparing methylation levels with reference cohorts, and using multidimensional analysis to classify the presence of tumors based on methylation patterns.

Benefits of technology

Enhances the sensitivity and specificity of tumor detection by accurately distinguishing between healthy individuals and those with tumors, reducing false positives and negatives, and providing a reliable method for early-stage tumor identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186258000029
    Figure 2025186258000029
  • Figure 2025186258000030
    Figure 2025186258000030
  • Figure 2025186258000031
    Figure 2025186258000031
Patent Text Reader

Abstract

To provide a method with sensitivity and / or specificity to detect a tumor at an early stage.SOLUTION: Cell-free DNA molecules in a mixture of a biological sample can be analyzed to detect viral DNA. Methylation of viral DNA molecules at one or more sites in the viral genome can be determined. Mixture methylation level (s) can be measured based on one or more amounts of the plurality of cell-free DNA molecules methylated at a set of site (s) of the particular viral genome. The mixture methylation level (s) can be determined in various ways, e.g., as a density of cell-free DNA molecules that are methylated at a site or across multiple sites or regions. The mixture methylation level (s) can be compared to reference methylation level (s), e.g., determined from at least two cohorts of other subjects. The cohorts can have different classifications (including the first condition) associated with the particular viral genome. A first classification of whether the subject has the first condition can be determined based on the comparing.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and is a non-provisional application of U.S. Provisional Patent Application No. 62 / 537,328, filed July 26, 2017, entitled "Enhancement Of Cancer Screening Using Cell-Free Viral Nucleic Acids," the entire contents of which are incorporated herein by reference for all purposes.

[0002] The discovery that tumor cells release tumor-derived DNA into the bloodstream has sparked the development of noninvasive methods that can determine the presence, location, and / or type of tumor in a subject using acellular samples (such as plasma). Many tumors can be treatable if detected early in their development. However, current methods may lack the sensitivity and / or specificity to detect tumors in their early stages and may return numerous false-positive or false-negative results. For example, certain viruses are associated with cancer, but viral DNA can be detected in subjects without cancer, resulting in false-positive results.

[0003] The sensitivity of a test may refer to the likelihood that a subject who tests positive for a condition will also be positive for that condition. The specificity of a test may refer to the likelihood that a subject who tests negative for a condition will also be negative for that condition. In assays for early tumor detection, issues of sensitivity and specificity may be exaggerated. For example, the samples on which such tumor detection methods are performed may contain relatively small amounts of tumor material, and the condition itself may have a relatively low prevalence among individuals tested at an early stage. Thus, there is a clinical need for methods with higher sensitivity and / or specificity for tumor detection. Summary of the Invention

[0004] Embodiments provide systems, devices, and methods for analyzing biological samples from subjects in the animal kingdom, such as humans. Cell-free DNA molecules in a mixture of biological samples can be analyzed to detect viral DNA, for example, by determining their location in a specific viral genome. The methylation status of viral DNA at one or more sites in the viral genome can be determined. The mixture methylation level(s) can be measured based on the amount of one or more of a plurality of cell-free DNA molecules methylated at a set of site(s) in a specific viral genome. The mixture methylation level(s) can be determined in various ways, for example, as the percentage / density of cell-free DNA molecules that are methylated at a specific site or across multiple sites, and optionally across multiple regions (each region containing one or more sites).

[0005] The mixed methylation level(s) can be compared with reference methylation level(s), for example, determined from at least two cohorts of other subjects. The cohorts can have different classifications (including a first pathology) associated with a specific viral genome. The other cohort(s) can correspond to other pathology(s). This comparison can be performed in various ways, for example, by forming multidimensional points of N methylation levels and determining the difference from the N reference methylation levels. A first classification of whether the subject has the first pathology can be determined based on this comparison.

[0006] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments relate to systems, devices, and computer-readable media associated with the methods described herein.

[0007] A better understanding of the nature and advantages of embodiments of the present disclosure may be obtained with reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0008] [Figure 1] 1 shows the design of a capture probe for targeted bisulfite sequencing according to an embodiment of the present disclosure. [Figure 2] 1 shows the methylation density of CpG sites across the Epstein-Barr virus (EBV) genome in patients with infectious mononucleosis, nasopharyngeal carcinoma (NPC), and natural killer (NK)-T cell lymphoma, according to embodiments of the present disclosure. [Figure 3] 1 shows the plasma EBV DNA methylation profile in a patient (AL038) from the screening cohort with early stage NPC (Stage I), according to an embodiment of the present disclosure. [Figure 4] 1 shows the difference in methylation density of CpG sites across the EBV genome between two patients with different disease states, according to an embodiment of the present disclosure. [Figure 5] 1 shows the difference in methylation density of CpG sites across the EBV genome between two patients with NPC (TBR1392 and TBR1416), according to an embodiment of the present disclosure. [Figure 6] 1 shows the difference in methylation patterns of plasma EBV DNA between a patient with early stage NPC (AO050) and a subject with a false-positive plasma EBV DNA result (HB002), according to an embodiment of the present disclosure. [Figure 7A-7C] 1 is a dot plot showing the methylation density of CpG sites across the EBV genome of one patient (x-axis) and the corresponding methylation density of the same CpG sites in another patient (y-axis), according to an embodiment of the present disclosure. [Figure 8] 1 shows the methylation percentage of plasma EBV DNA based on CpG sites of the EBV genome in subjects with infectious mononucleosis (IM) (n=2), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6) according to an embodiment of the present disclosure. [Figure 9] 1 illustrates mining of differentially methylated regions (DMRs) that meet a first selection criterion, according to an embodiment of the present disclosure. [Figure 10]FIG. 10 is a table listing the genomic coordinates of differentially methylated regions that meet the criteria set forth in FIG. 9. [Figure 11] 1 shows the methylation percentage of plasma EBV DNA based on 821 CpG sites within the 39 DMRs described in FIG. 10 in subjects with infectious mononucleosis (IM) (n=2), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6), according to an embodiment of the present disclosure. [Figure 12] 1 illustrates mining of differentially methylated regions (DMRs) that meet a second selection criterion, according to an embodiment of the present disclosure. [Figure 13] 12 shows the methylation percentage of plasma EBV DNA based on the 46 DMRs identified in FIG. 12 in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to an embodiment of the present disclosure. [Figure 14] 10 illustrates mining of representative methylation consensus regions that meet the third selection criterion, according to an embodiment of the present disclosure. [Figure 15] 1 shows the methylation percentage of plasma EBV DNA based on the "representative" CpG sites described in FIG. 12 in the same groups of subjects with infectious mononucleosis (IM) (n=2), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6), according to an embodiment of the present disclosure. [Figure 16] 1 shows examples of CpG sites with methylation percentages across more than 80% of sites in pooled sequencing data of three cases with persistently positive plasma EBV DNA and a mean value of less than 20% of sites in three subjects with NPC, according to embodiments of the present disclosure. [Figure 17] 1 shows examples of CpG sites with methylation percentages across less than 20% of sites in pooled sequencing data of three cases with persistently positive plasma EBV DNA and more than 80% of sites in three subjects with NPC, according to embodiments of the present disclosure. [Figure 18]1 shows cluster dendrograms using hierarchical clustering analysis based on methylation pattern analysis of plasma EBV DNA for six NPC patients (including four patients with early stage disease from the screening cohort of the present invention), two patients with extranodal NK-T cell lymphoma, and two patients with infectious mononucleosis, according to embodiments of the present disclosure. [Figure 19] 1 shows a cluster dendrogram using hierarchical clustering analysis based on methylation pattern analysis of plasma EBV DNA for six NPC patients (including four early-stage NPC patients from the screening cohort of the present invention) and three non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. [Figure 20] Heatmap 2000 showing the methylation levels of all non-overlapping 500 bp regions in the whole EBV genome in patients with nasopharyngeal carcinoma, NK-T cell lymphoma, and infectious mononucleosis. [Figure 21] 1 shows size profiles of the size distribution of sequenced plasma DNA fragments mapped to the EBV genome and human genome in two NPC patients (TBR1392 and TBR1416) and two infectious mononucleosis patients (TBR1610 and TBR1661), as well as three non-NPC subjects (AF091, HB002, and HF020) with persistently positive plasma EBV DNA on serial analyses, according to embodiments of the present disclosure. [Figure 22] 1 shows size ratios in six NPC patients and three subjects persistently positive for plasma EBV DNA, according to embodiments of the present disclosure. [Figure 23] 1 shows EBV DNA size ratios in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to embodiments of the present disclosure. [Figure 24]1 shows the percentage of plasma EBV DNA reads (plasma DNA reads mapped to the EBV genome) among all plasma DNA reads sequenced in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to embodiments of the present disclosure. [Figure 25] 1 is a plot of plasma EBV DNA read percentages and corresponding size ratio values ​​for NPC patients, transiently positive, and non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. [Figure 26] 1 is a plot of plasma EBV DNA read percentages and corresponding methylation percentage values ​​for NPC patients, transiently positive, and non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. [Figures 27A-27B] 1 shows a three-dimensional plot of plasma EBV DNA read percentages and corresponding size ratio and methylation percentage values ​​for NPC patients, transiently positive, and non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. [Figures 28A-28B] 1 shows receiver operating characteristic (ROC) curve analysis of various combinations of number-based, size-based, and methylation-based analysis according to an embodiment of the present disclosure. [Figure 29] Clinical staging of five cases of HPV-positive head and neck squamous cell carcinoma (HPV+ve HNSCC) is shown. [Figure 30] 1 shows plasma HPV DNA methylation profiles in individual patients with HPV-positive head and neck squamous cell carcinoma (HPV+ve HNSCC), according to embodiments of the present disclosure. [Figure 31] 1 shows the methylation levels of all CpG sites in the entire HPV genome in two patients with HPV+ve HNSCC, according to an embodiment of the present disclosure. [Figure 32A-32B]1 shows the percentage of hepatitis B virus (HBV) DNA reads (plasma DNA reads mapped to the HBV genome) and the methylation percentage of all CpG sites in the entire HBV genome for nine patients with chronic hepatitis B virus infection (HBV) and ten patients with hepatocellular carcinoma (HCC), according to an embodiment of the present disclosure. [Figure 33] 1 is a flow chart illustrating a method of analyzing a biological sample from an animal subject to determine a classification of a first disease state, according to an embodiment of the present disclosure. [Figure 34] 1 illustrates a system according to one embodiment of the present invention. [Figure 35] 1 shows a block diagram of an exemplary computer system usable with systems and methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Appendix A shows a list of individual CpG sites in the entire EBV genome with differential methylation levels where the difference in methylation percentage at these CpG sites between the pooled sequence data of three subjects with persistently positive EBV DNA and three NPC patients was greater than 20%. Sites marked with * represent a difference in methylation percentage of greater than 40%, ** greater than 60%, and *** greater than 80%.

[0010] term The terms "sample," "biological sample," or "patient sample" are meant to include any tissue or material derived from a living or dead subject. A biological sample can be an acellular sample, which can contain a mixture of nucleic acid molecules from a subject and nucleic acid molecules from a pathogen, such as a virus in some cases. A biological sample generally contains nucleic acid (e.g., DNA or RNA) or fragments thereof. The term "nucleic acid" can generally refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. The nucleic acid in a sample can be acellular nucleic acid. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A biological sample can be a bodily fluid such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., testicular), vaginal washing, pleural effusion, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body, etc. A stool sample can also be used. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free). The centrifugation protocol can include, for example, obtaining a fluid portion at 3,000 g for 10 minutes, followed by recentrifugation at 30,000 g for an additional 10 minutes to remove residual cells.

[0011] As used herein, the term "fragment" (e.g., DNA fragment) may refer to a portion of a polynucleotide or polypeptide sequence comprising at least three consecutive nucleotides. Nucleic acid fragments can retain biological activity and / or some characteristics of the parent polypeptide. Nucleic acid fragments can be double-stranded or single-stranded, methylated or unmethylated, intact or nicked, and complexed or uncomplexed with other macromolecules, such as lipid particles or proteins. In one example, nasopharyngeal carcinoma cells can release Epstein-Barr virus (EBV) DNA fragments into the bloodstream of a subject, e.g., a patient. These fragments can contain one or more BamHI-W sequence fragments, which can be used to detect levels of tumor-derived DNA in plasma. BamHI-W sequence fragments correspond to sequences that can be recognized and / or digested using the BamHI restriction enzyme. The BamHI-W sequence may refer to the sequence 5'-GGATCC-3'.

[0012] Tumor-derived nucleic acid can refer to any nucleic acid released from tumor cells, including pathogen nucleic acid from pathogens within tumor cells. For example, Epstein-Barr virus (EBV) DNA can be released from cancer cells in subjects with nasopharyngeal carcinoma (NPC).

[0013] The term "assay" generally refers to a technique for determining a characteristic of a nucleic acid. An assay (e.g., a first assay or a second assay) generally refers to a technique for determining the amount of a nucleic acid in a sample, the genomic identity of a nucleic acid in a sample, the copy number variation of a nucleic acid in a sample, the methylation state of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation state of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the nucleic acid characteristics mentioned herein. Nucleic acid characteristics include sequence, amount, genomic identity, copy number, methylation state at one or more nucleotide positions, nucleic acid size, nucleic acid mutation at one or more nucleotide positions, and nucleic acid fragmentation pattern (e.g., the nucleotide position(s) at which the nucleic acid fragments). The term "assay" may be used interchangeably with the term "method." An assay or method has a particular sensitivity and / or specificity, and its relative usefulness as a diagnostic tool can be measured using the ROC-AUC statistic.

[0014] The term "random sequencing" as used herein generally refers to sequencing in which the nucleic acid fragments to be sequenced are not specifically identified or predetermined before the sequencing procedure. No sequence-specific primers are required to target specific gene loci. In some embodiments, adapters are added to the ends of the fragments, and sequencing primers are bound to the adapters. Therefore, any fragment can be sequenced with the same primer that binds to the same universal adapter, and thus the sequencing can be random. Random sequencing can also be used to perform massively parallel sequencing.

[0015] A "sequence read" generally refers to a chain of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short chain of nucleotides (e.g., about 20-150 bases) sequenced from a nucleic acid fragment, a short chain of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained, for example, using sequencing techniques or probes in a variety of ways, such as with hybridization arrays or capture probes, or with amplification techniques such as polymerase chain reaction (PCR) or linear amplification using single primers or isothermal amplification, or based on biophysical measurements such as mass spectrometry.

[0016] A "methylome" provides a measure of the amount of DNA methylation at multiple sites or loci in a genome (e.g., a human or other animal genome, or a viral genome). The methylome can correspond to the entire genome, a substantial portion of the genome, or a relatively small portion(s) of the genome. Examples of methylomes of interest include the methylomes of tumor cells (e.g., nasopharyngeal carcinoma, hepatocellular carcinoma, cervical carcinoma), viral methylomes (e.g., EBV present in a subject's healthy or tumor cells), bacterial methylomes, and organs (e.g., brain cells, bones, lungs, heart, muscles, and kidneys) that can contribute DNA to bodily fluids (e.g., plasma, serum, sweat, saliva, urine, reproductive secretions, semen, fecal fluid, diarrheal fluid, cerebrospinal fluid, gastrointestinal secretions, ascites, pleural fluid, intraocular fluid, fluid from edema (e.g., testes), cyst fluid, pancreatic secretions, intestinal secretions, sputum, tears, aspirates from the breast and thyroid, etc.). Organs may also be transplanted organs. The methylome of a fetus is another example.

[0017] A "plasma methylome" is a methylome determined from the plasma or serum of an animal (e.g., a human). The plasma methylome is an example of a cell-free methylome because plasma and serum contain cell-free DNA. The plasma methylome is also an example of a mixed methylome because it is a mixture of fetal / maternal methylomes, tumor / patient methylomes, DNA from different tissues or organs, donor / recipient methylomes in organ transplant situations, and / or DNA from different genomes (e.g., animal genomes and bacterial / viral genomes).

[0018] A "site" (also called a "genomic site") corresponds to a single site, which can be a single base position, or a group of correlated base positions, e.g., a CpG site, or a larger group of correlated base positions. A "locus" can correspond to a region that includes multiple sites. A locus can contain only one site, which would make the locus equivalent to the site in its context.

[0019] A "methylation index" for each genomic site (e.g., a CpG site) can refer to the proportion of DNA fragments (e.g., as determined from sequence reads or probes) that exhibit methylation at that site across the total number of reads covering that site. A "read" can correspond to information obtained from a DNA fragment (e.g., the methylation state of the site). Reads can be obtained using reagents (e.g., primers or probes) that preferentially hybridize to DNA fragments of a particular methylation state. Typically, such reagents are applied after treatment with a process that specifically modifies or specifically recognizes DNA molecules depending on their methylation state, such as bisulfite conversion, or a methylation-sensitive restriction enzyme, or a methylation-binding protein, or an anti-methylcytosine antibody. In another embodiment, single-molecule sequencing techniques that recognize methylcytosine and hydroxymethylcytosine can be used to elucidate the methylation state and determine the methylation index.

[0020] The "methylation density" of a region can refer to the number of reads at a site within the region that show methylation divided by the total number of reads that cover the site in this region. This site can have specific characteristics, for example, it can be a CpG site. Thus, the "CpG methylation density" of a region refers to the number of reads that show CpG methylation divided by the total number of reads that cover the CpG site in this region (e.g., a specific CpG site, a CpG site within a CpG island, or a larger region). For example, the methylation density of each 100 kb bin in the human genome can be determined from the total number of unconverted cytosines (corresponding to methylated cytosines) at CpG sites after bisulfite treatment as a percentage of all CpG sites covered by sequence reads mapped to the 100 kb region. This analysis can also be performed for other bin sizes, such as 500 bp, 5 kb, 10 kb, 50 kb, or 1 Mb. A region can be the whole genome, or a chromosome, or a part of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site is the same as the methylation density of that region if the region contains only that CpG site. "Percentage of methylated cytosines" can refer to the total number of analyzed cytosine residues in this region, i.e., the number of cytosine sites "C" that are shown to be methylated (e.g., unconverted after bisulfite conversion), including cytosines outside the context of CpG. Methylation index, methylation density, and percentage of methylated cytosines are examples of "methylation levels," which can include other ratios, including the number of methylated reads at a site.Aside from bisulfite conversion, other processes known to those skilled in the art can be used to examine the methylation state of DNA molecules, including, but not limited to, by methylation-state-sensitive enzymes (e.g., methylation-sensitive restriction enzymes), methylation-binding proteins, single-molecule sequencing using methylation-state-sensitive platforms (e.g., nanopore sequencing (Schreiber et al. Proc Natl Acad Sci 2013;110:18910-18915), and Pacific Biosciences single-molecule real-time analysis (Flusberg et al. Nat Methods 2010;7:461-465)).

[0021] A "methylation profile" (also called a methylation status) contains information related to DNA methylation for a region. Information related to DNA methylation can include, but is not limited to, the methylation index of CpG sites, the methylation density of CpG sites in a region, the distribution of CpG sites across a contiguous region, the methylation pattern or level of each individual CpG site within a region containing two or more CpG sites, and non-CpG methylation. A methylation profile of a substantial portion of the genome (e.g., covering more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%) can be considered equivalent to a methylome. "DNA methylation" in mammalian genomes typically refers to the addition of a methyl group to the 5' carbon of cytosine residues in CpG dinucleotides (i.e., 5-methylcytosine). DNA methylation can occur at cytosine in other contexts, such as CHG and CHH, where H is adenine, cytosine, or thymine. Methylation of cytosine can also occur in the form of 5-hydroxymethylcytosine. 6 Non-cytosine methylations such as -methyladenine have also been reported.

[0022] "Methylation-recognition sequence" refers to a sequencing method that can ascertain the methylation status of a DNA molecule during the sequencing process, including, but not limited to, bisulfite sequencing, or methylation-sensitive restriction enzyme digestion, immunoprecipitation using anti-methylcytosine antibodies or methylation-binding proteins, or single-molecule sequencing that allows for elucidation of the methylation status. "Methylation-recognition assay" or "methylation-sensitive assay" can include both sequencing and non-sequencing-based methods, such as MSP, probe-based interrogation, hybridization, restriction enzyme digestion followed by densitometry, anti-methylcytosine immunoassay, mass spectrometry interrogation of the proportion of methylated cytosine or hydroxymethylcytosine, and immunoprecipitation without sequencing.

[0023] A "tissue" corresponds to a group of cells that group together as a functional unit. Two or more types of cells can be found within a single tissue. Different types of tissue can consist of different cell types (e.g., liver cells, lung cells, or blood cells), but can also correspond to tissues from different organisms (host vs. virus) or healthy vs. tumor cells. The term "tissue" can generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the terms "tissue" or "tissue type" can be used to refer to the tissue from which cell-free nucleic acid is derived. In one example, viral nucleic acid fragments can be derived from blood tissue, e.g., Epstein-Barr virus (EBV). In another example, viral nucleic acid fragments can be derived from tumor tissue, e.g., EBV or human papillomavirus (HPV) infection.

[0024] A "separation value" (or relative abundance) corresponds to the difference or ratio between two values, such as two DNA molecular weights, two contribution ratios, or two methylation levels (such as a sample (mixture) methylation level and a reference methylation level). A separation value can be a simple difference or ratio. For example, the direct ratio of x / y is a separation value, as is x / (x+y). A separation value can include other factors, such as multiplicative factors. As another example, the difference or ratio of a function of values, such as the difference or ratio of the natural logarithms (ln) of two values, can be used. A separation value can include differences and / or ratios. A methylation level is an example of relative abundance, such as the relative abundance of a methylated DNA molecule (e.g., at a particular site) relative to other DNA molecules (e.g., all other DNA molecules at a particular site or unmethylated DNA molecules). The molecular weight of other DNA can serve as a normalization factor. As another example, the intensity (e.g., fluorescence or field intensity) of a methylated DNA molecule relative to the intensity of all or unmethylated DNA molecules can be determined. Relative abundance can also include intensity per volume.

[0025] As used herein, the term "classification" refers to any number(s) or other feature(s) associated with a particular characteristic of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a particular level of pathology (e.g., cancer). Classification may be binary (e.g., positive or negative) or may have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1).

[0026] The terms "cutoff," "threshold," or reference level may refer to a predetermined number used in an operation. The threshold or reference value may be a value above or below which a particular classification is applied, for example, a classification of a condition, such as whether a subject has a condition or the severity of the condition. The cutoff may be predetermined with or without reference to sample or subject characteristics. For example, the cutoff may be selected based on the age or sex of the subject being tested. The cutoff may be selected after and based on the output of test data. For example, a particular cutoff may be used when the sequencing of a sample reaches a certain depth. As another example, a reference subject with a known classification of one or more conditions and a measured characteristic value (e.g., methylation level) may be used to determine a reference level that distinguishes between different conditions and / or classifications of conditions (e.g., whether a subject has a condition). Any of these terms may be used in any of these contexts.

[0027] The terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" can be used interchangeably to generally describe a sample that does not have a specific pathology or is otherwise healthy. In one example, the methods disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the subject's healthy tissue. In another example, the reference sample is a sample taken from a subject with a disease, such as cancer or a specific stage of cancer. The reference sample can be obtained from the subject or a database. The reference generally refers to a reference genome used to map sequence reads obtained from sequencing a sample from a subject. The reference genome generally refers to a haploid or diploid genome to which sequence reads from a biological sample and a natural sample can be aligned and compared. For a haploid genome, only one nucleotide is present at each locus. For a diploid genome, heterozygous loci can be identified; such loci have two alleles, and either allele can allow alignment to the locus. The reference genome can correspond to a virus, for example, by including one or more viral genomes.

[0028] As used herein, the phrase "healthy" generally refers to a subject who is in good health. Such a subject exhibits the absence of malignant or non-malignant disease. A "healthy individual" may have other diseases or conditions unrelated to the condition being assayed that are not normally considered "healthy."

[0029] The terms "cancer" and "tumor" are used interchangeably and generally refer to an abnormal mass of tissue whose growth exceeds and is uncoordinated with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are generally well differentiated, characteristically grow slower than malignant tumors, and remain localized at the site of origin. Furthermore, benign tumors lack the ability to infiltrate, invasive, or metastasize to distant sites. "Malignant" tumors are generally poorly differentiated (anaplastic) and exhibit characteristically rapid growth accompanied by progressive infiltration, invasion, and destruction of surrounding tissue. Furthermore, malignant tumors have the ability to metastasize to distant sites. "Stage" can be used to describe the progression of malignant tumors. Early-stage cancers or malignant tumors have a smaller tumor burden in the body than later-stage malignant tumors and are generally associated with fewer symptoms, a better prognosis, and better treatment outcomes. Late or advanced stage cancers or malignant tumors are often associated with distant metastasis and / or lymphatic spread.

[0030] The term "level of cancer" (or more generally, "level of disease" or "level of pathology") can refer to whether cancer is present (i.e., present or absent), the stage of cancer, tumor size, whether there is metastasis, total tumor burden in the body, the response of cancer to treatment, and / or other measures of cancer severity (such as cancer recurrence). The level of cancer can be a number or other indicia, such as a symbol, alphabetic letter, and color. The level can be zero. The level of cancer can also include premalignant or precancerous conditions. The level of cancer can be used in various ways. For example, screening can check for the presence of cancer in a person who was not previously known to have cancer. Evaluation can be performed on a person who has been diagnosed with cancer to monitor the progression of the cancer over time, study the effectiveness of therapy, or determine a prognosis. In one embodiment, the prognosis can be expressed as the likelihood that the patient will die from the cancer, or the likelihood that the cancer will progress for a certain duration or after a certain time, or the likelihood that the cancer will metastasize. Detection can mean "screening" or checking whether a person with suggestive features of cancer (e.g., symptoms or other positive tests) has cancer. "Level of pathology" can refer to the level of pathology associated with a pathogen, which may be as described above for cancer. The level of disease / pathology may also be as described above for cancer. If cancer is associated with a pathogen, the level of cancer can be a type of level of pathology.

[0031] The terms "size profile" and "size distribution" generally relate to the size of DNA fragments in a biological sample. A size profile can be a histogram that provides the distribution of a certain amount of DNA fragments of various sizes. Various statistical parameters (also called size parameters or simply parameters) can distinguish one size profile from another. One parameter is the proportion of DNA fragments of a particular size or size range relative to all DNA fragments or DNA fragments of other sizes or ranges.

[0032] The term "false positive" (FP) can refer to a subject who does not have a pathological condition. A false positive generally refers to a subject who does not have a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, a non-malignant disease, or who is otherwise healthy. The term false positive generally refers to a subject who does not have a pathological condition but is identified as having a pathological condition by an assay or method of the present disclosure.

[0033] The term "sensitivity" or "true positive rate" (TPR) can refer to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to accurately identify the proportion of a population that truly has a disease state. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to accurately identify one or more markers that indicate cancer.

[0034] The term "specificity" or "true negative rate" (TNR) can refer to the number of true negatives divided by the sum of the number of true negatives and false positives.Specificity can characterize the ability of an assay or method to accurately identify the proportion of a population that is truly free of a disease state.For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that are free of cancer.In another example, specificity can characterize the ability of a method to correctly identify one or more markers that indicate cancer.

[0035] The term "ROC" or "ROC curve" may refer to a receiver operating characteristic curve. ROC curves can graphically represent the performance of a binary classification system. For any given method, a ROC curve can be generated by plotting sensitivity against specificity at various threshold settings. The sensitivity and specificity of a method for detecting the presence of a tumor in a subject can be determined at various concentrations of tumor-derived nucleic acid in the subject's plasma sample. Furthermore, the value or expected value of any unknown parameter can be determined using at least one of the three obtained parameters (sensitivity, specificity, threshold setting, etc.) and the ROC curve. The unknown parameter can be determined using a curve fit to the ROC curve. The term "AUC" or "ROC-AUC" generally refers to the area under the receiver operating characteristic curve. This metric can provide a measure of the diagnostic utility of a method, taking into account both the sensitivity and specificity of the method. Generally, the ROC-AUC ranges from 0.5 to 1.0, with values ​​closer to 0.5 indicating a method with limited diagnostic utility (e.g., low sensitivity and / or low specificity) and values ​​closer to 1.0 indicating a method with high diagnostic utility (e.g., high sensitivity and / or high specificity). See, e.g., Pepe et al., "Limitations of the Odds Ratio in Gauging the Performance of a Diagnostic, Prognostic, or Screening Marker," Am. J. Epidemiol. 2004, 159(9):882-890, incorporated herein by reference. Additional approaches to characterizing diagnostic utility using likelihood functions, odds ratios, information theory, predictive value, calibration (including goodness of fit), and reclassification measures are summarized in Cook, "Use and Misuse of the Receiver Operating Characteristic Curve in Risk Prediction," Circulation 2007, 115:928-935, which is incorporated herein by reference in its entirety.

[0036] The term "about" or "approximately" can mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which depends in part on the method of measuring or determining the value, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations, as is customary in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, or more preferably within 2-fold of a value. When a specific value is described in this application and claims, unless otherwise specified, the term "about" should be assumed to be within an acceptable error range of the particular value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.

[0037] The terms used herein are for the purpose of describing particular cases only and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The use of "or" is intended to mean "inclusive or" rather than "exclusive or," unless specifically stated to the contrary. The term "based on" is intended to mean "based at least in part on." Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0038] This disclosure describes an approach to distinguish between different EBV-associated diseases, malignancies, conditions, or completely healthy individuals based on the analysis of methylation patterns of circulating EBV DNA fragments in the blood. Analysis of the methylation patterns of cell-free EBV DNA molecules has several applications and practical uses. The feasibility of analyzing the methylation of cell-free viral molecules noninvasively will enhance clinical applications in the context of screening, predictive medicine, risk stratification, surveillance, and prognosis.

[0039] The embodiments can distinguish between subjects with different virus-related pathologies (e.g., NPC patients) and apparently healthy subjects with detectable plasma EBV DNA, even in a single-time point analysis from a single blood draw. The embodiments can also be used for screening or detecting whether a subject has disease or cancer, for disease monitoring in cancer patients, for prognosis, and for disease or cancer risk prediction (i.e., for predicting whether a subject will develop disease or cancer in the future). This approach can be generalized to viruses other than EBV. Thus, this approach is a general approach for identifying viral DNA-based biomarkers. I. Cancer and Viruses

[0040] Both DNA and RNA viruses have been shown to be capable of causing cancer in humans. In some embodiments, the subject may have cancer caused by a virus (e.g., an oncovirus). In some embodiments, the subject may have cancer, and the cancer may be detectable using viral DNA. In RNA analysis, nucleic acids exist as complementary DNA (cDNA), which can be copied from RNA and serve as a medium for replication in host cells. These cDNAs have methylation and may be used in embodiments.

[0041] Various viral infections are associated with various cancers and other pathologies. For example, EBV infection is closely associated with NPC and natural killer (NK) T-cell lymphoma, Hodgkin's lymphoma, gastric cancer, and infectious mononucleosis. Hepatitis B virus (HBV) and hepatitis C virus (HCV) infections are associated with an increased risk of developing hepatocellular carcinoma (HCC). Human papillomavirus (HPV) infections are associated with an increased risk of developing cervical cancer (CC) and head and neck squamous cell carcinoma (HNSCC). While the examples focus on EBV, the techniques are equally applicable to cancers and other pathologies involving HPV, HBV, and other viruses, particularly those associated with cancer. A.EBV

[0042] It is estimated that 95% of the world's population has lifelong, asymptomatic Epstein-Barr virus (EBV) infection, which allows the virus to persist in memory B cells of healthy individuals (Young et al. Nat Rev Cancer 2016 16(12):789-802). A small number of subjects develop symptomatic infection, manifesting as infectious mononucleosis due to viral infection. EBV is also considered an oncogenic virus in association with several malignancies or cancer-like syndromes of epithelial and hematologic origin, including nasopharyngeal carcinoma (NPC), gastric cancer, Burkitt lymphoma, Hodgkin lymphoma, natural killer T-cell (NK-T) lymphoma, and post-transplant lymphoproliferative disorder (PTLD).

[0043] The diagnostic and prognostic role of circulating EBV DNA has been explored in patients with EBV-associated malignancies. In this regard, plasma EBV DNA has been established as a biomarker for NPC (Lo et al. Cancer Res 1999;59:1188-91). Periodic surveillance using plasma EBV DNA is recommended for detecting residual disease and recurrence in patients with a confirmed diagnosis of NPC (Lo et al. Cancer Res 1999;59:5452-5, Chan et al. J Natl Cancer Inst 2002;94:1614-9, Leung et al. Cancer 2003,98(2),288-91, and Leung et al. Ann Oncol 2014;25(6):1204-8). Plasma EBV DNA has also been shown to have prognostic significance in other EBV-associated malignancies, including Hodgkin lymphoma (Kanakry et al. Blood 2013;121(18):3547-3553), extranodal NK-T cell lymphoma (Wang et al. Oncotarget 2015;6(30):30317-26, Kwong et al. Leukemia 2014;28(4):865-870), and PTLD (Gulley and Tang. Clin Microbiol Rev 2010;23(2):350-66).

[0044] However, not all subjects with such infections develop the associated cancer. The source of plasma EBV DNA must be different in individuals without NPC. Unlike the sustained release of EBV DNA from NPC cells into the circulation, the supply of EBV DNA in individuals without NPC only provides such DNA transiently. B. False positive

[0045] In the context of cancer screening, we recently conducted a large-scale prospective study on NPC screening using plasma EBV DNA analysis by quantitative PCR (qPCR) (Chan et al. N Engl J Med 2017;377:513-522). We analyzed plasma EBV DNA levels in all recruited subjects (screening cohort) who were asymptomatic for NPC at enrollment. Subjects with detectable plasma EBV DNA were retested for EBV DNA 4 weeks after the initial test. Of the 20,174 recruited subjects, 1,112 had detectable plasma EBV DNA at the initial test. 309 subjects remained positive in follow-up tests based on plasma EBV DNA measurement. Thirty-four subjects with persistently positive plasma EBV DNA results were subsequently confirmed to have NPC by endoscopy and magnetic resonance imaging (MRI). As previously mentioned, plasma EBV DNA could be detected in apparently healthy individuals without NPC or other EBV-associated malignancies.

[0046] In 20,174 subjects screened for NPC, the false-positive rate for plasma EBV DNA based on a single-timepoint analysis was approximately 5% ((1112-34) / (20174-34) = 5.3%). Two serial EBV DNA analyses reduced the false-positive rate to 1.5%. However, serial plasma EBV DNA testing requires the collection of additional blood samples from subjects with an initial positive result, which can pose logistical challenges. Furthermore, a significant proportion of subjects with positive plasma EBV DNA results do not have NPC (96% of subjects with positive results in a single-timepoint analysis, determined as (1112-34) / 1112, do not have NPC). Subjects with false-positive results would require serial evaluations and unnecessary investigations, such as endoscopy and MRI, for a definitive diagnosis. All of this can lead to patient anxiety and high follow-up costs. Therefore, we aim to distinguish NPC patients from subjects with false-positive plasma EBV DNA results using a single-timepoint blood analysis. In this instance, the false plasma EBV DNA positivity rate is considered the non-NPC positivity rate, or also called the one-time positivity rate. C. Use of Methylation

[0047] Previous studies have reported different types of viral latency (types 0, 1, 2, and 3), which are defined by the latency-associated viral gene transcription patterns found in different EBV-associated malignancies (Young et al. Nat Rev Cancer 2016;16(12):789-802). Viral latency is defined by the latency-associated gene transcription patterns. Therefore, viruses with different types of viral latency have different patterns of viral gene transcription. Different EBV-associated diseases or pathologies with the same type of viral latency can have similar viral gene transcription patterns.

[0048] Among different latency types, there are distinct viral gene expression profiles and distinct methylation states of different viral gene promoters, including the replication origin, C promoter, W promoter, Q promoter, and LMP1 / 2 promoter (Woeller et al. Curr Opin Virol 2013;3(3):260-5). DNA methylation contributes to the regulation of gene expression, and it has been suggested that there are latency-type-specific methylation patterns (Lieberman. Nat Rev Microbiol 2013;11(12):863-75). In one example, a previous study found a methylation state of the C promoter that was compatible with the latency type II-specific methylation pattern in EBV DNA derived from nasopharyngeal brush cytology samples of NPC patients using methylation-specific PCR (Ramayanti et al. Int J Cancer 140,149-162). However, different EBV-associated diseases or conditions can have the same type of viral latency and therefore similar viral gene transcription patterns (examples are described in the next paragraph). Therefore, viral latency does not correlate with the stage of the disease or cancer.

[0049] Different EBV-associated diseases with the same type of viral latency are expected to have similar methylation patterns (Tempera et al. Semin Cancer Biol 2014;26:22-9, Fejer et al. J Gen Virol 2008;89:1364-70). In one example, previous studies using methylation-specific PCR demonstrated similar methylation patterns across the EBV viral promoter region (exhibiting latency type I) in both B cells from healthy EBV-seropositive individuals and tumor tissue from EBV-positive lymphomas (Paulson et al. J Virol 1999;73:9959-68).

[0050] Previous studies have attempted to study the methylation profile of EBV by amplicon sequencing of bisulfite-converted DNA from cell lines and tissue samples from different EBV-associated diseases (Fernandex et al. Genome Res 2009;19(3):438-51). The 77 amplicons designed covered the transcription start sites of 94 different EBV latent and lytic genes, as well as two structural RNAs (EBER1 and EBER2). The methylation status (either methylated or unmethylated) of transcription start sites across the entire EBV genome was assessed. These results, in contrast to quantification, showed that free viral DNA lacked DNA methylation, and only numerous methylated EBV transcription start sites were present in viral DNA derived from cell lines or tissue samples from EBV-associated malignancies. Importantly, samples from different malignancies (i.e., NPC and different lymphomas) clustered together based on the methylation patterns of transcription start sites, and no transiently or persistently positive subjects were identified. Based on their methylation patterns, different malignant conditions could not be distinguished.

[0051] Most previous studies have focused on analyzing viral methylation profiles in tumor and cell line samples. These tumor samples need to be obtained through invasive procedures such as surgical biopsy, which may limit diagnostic applications such as screening and serial monitoring. Furthermore, previous studies have focused on qualitative aspects rather than quantitative results.

[0052] Despite the above reported data, we investigate the feasibility of distinguishing between different EBV-associated diseases exhibiting the same type of viral latency. In contrast to the above reported data, we describe a method based on the analysis of the methylation profile of plasma EBV DNA sequences that can distinguish between different EBV-associated diseases or disease stages. For example, instead of analyzing only the methylation status (methylated or unmethylated) of viral gene promoters, we investigated the methylation level of each CpG site in cell-free EBV DNA molecules in a genome-wide manner with high resolution. Surprisingly, our data reveal that methylation analysis of cell-free EBV DNA molecules can distinguish between different EBV-associated pathologies and malignancies with the same latency type. Thus, our data provide new information about cell-free EBV DNA methylation patterns that goes beyond the inherent variability of the latency type.

[0053] Embodiments can analyze the methylation patterns of cell-free EBV DNA molecules in blood (e.g., plasma or serum). Embodiments of the present disclosure can also be used with other bodily fluids containing cell-free EBV DNA molecules, such as urine (Chan et al. Clin Cancer Res 2008;14(15):4809-13), serum, vaginal fluid, uterine or vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, etc. Fecal samples can also be used. A technical challenge is the low abundance and fragmented nature of viral molecules compared to analyzing tumor DNA in tissue samples. This disclosure demonstrates the feasibility of methylation analysis of cell-free viral molecules in a noninvasive manner. II. Measurement of methylation of cell-free EBV DNA molecules

[0054] Methylation level(s) can be measured at various sites in, for example, animal (such as human), viral, or other genomes. Methylation levels can be determined using methylation information for one or more sites, for example, CpG sites. Methylation information can include the number of methylated DNA molecules at a particular site, or an intensity signal corresponding to the amount of methylated / unmethylated DNA molecules. Methylation levels can provide the relative abundance between methylated and unmethylated DNA molecules; for example, the total DNA molecular weight or the unmethylated DNA molecular weight at a site can serve as a normalization factor.

[0055] For viral genomes, the average methylated CpG density (also called methylation density, MD) of a particular locus across the viral genome in plasma can be calculated using the following equation:

[0056]

number

[0057] where M is the number of methylated viral reads, and U is the number of unmethylated viral reads at CpG sites within a locus in the entire viral genome. If there are two or more CpG sites within a locus, M and U correspond to the number of methylated and unmethylated reads across the site, respectively. For example, the number of individual methylated or unmethylated DNA fragments can be determined using sequencing or digital PCR. As another example, rather than counting the number of specific reads, methylation density can be determined using real-time PCR to obtain a signal intensity ratio (e.g., the ratio of methylated intensity to unmethylated intensity). Therefore, if the intensity signal corresponds to multiple nucleic acids, analysis of the nucleic acids can be performed collectively. The specific format of the methylation level can vary, for example, the above-mentioned ratio or the ratio of M to U. A. Various Techniques for Assessing Methylation Levels

[0058] Different approaches can be used to determine methylation levels, for example, to determine the methylation profile across all or a substantial portion of a genome (e.g., a human genome or a viral genome). To comprehensively investigate methylation profiles, example embodiments can use massively parallel sequencing (MPS) of bisulfite-converted DNA to provide genome-wide information and quantitative assessment of methylation levels per nucleotide and per allele. Any methylation-sensitive assay can be used to determine the methylation levels of selected CpG sites. Other exemplary techniques include single molecule sequencing (e.g., nanopore sequencing (Simpson et al. Nat Methods 2017;14(4):407-410)), methylation-specific PCR (Herman et al. Proc Natl Acad Sci USA 1996;93(18):9821-9826), enzymes that specifically modify DNA molecules based on their methylation status (e.g., methylation-sensitive restriction enzymes), treatment with methylation-binding proteins (e.g., antibodies), or mass spectrometry-based methods (e.g., Lin et al. Anal Chem 2016;88(2):1083-7).

[0059] Various types of methylation can be analyzed. In some embodiments, 5-methylation of cytosine residues is used as an example. Other types of DNA methylation changes, such as hydroxymethylation or adenine methylation, can also be used. Therefore, techniques for detecting hydroxymethylation can also be used, such as oxidized bisulfite sequencing (Booth et al. Science 2012;336(6083):934-7) and tet-assisted bisulfite sequencing (NAT Protoc 2012;7(12):2159-70). Further details regarding the determination and use of methylation profiles can be found in U.S. Patent Publication Nos. 2015 / 0011403, 2016 / 0017419, and 2017 / 0029900, which are incorporated by reference in their entirety.

[0060] During bisulfite modification, unmethylated cytosines are converted to uracil and then to thymine, whereas methylated cytosines remain intact after PCR amplification (Frommer M, et al. Proc Natl Acad Sci USA 1992;89:1827-31). After sequencing and alignment, the methylation of individual CpG sites can be inferred from the number of methylated sequence reads (M) (methylated) and the number of unmethylated sequence reads (U) (unmethylated) at the cytosine residues of the CpG site. Bisulfite sequencing can be used to construct viral methylomes from the plasma of subjects with different virus-associated pathologies.

[0061] As described above, methylation profiling can be performed using massively parallel sequencing (MPS) of bisulfite-converted DNA. MPS of bisulfite-converted DNA can be performed in a random or shotgun manner, or in a targeted manner. For example, a region(s) of interest in bisulfite-converted DNA can be captured using a liquid-phase or solid-phase hybridization-based process, followed by MPS.

[0062] MPS can be performed using sequencing by synthesis platforms (e.g., Illumina HiSeq, NextSeq, NovaSeq platforms), sequencing by ligation platforms (e.g., Life Technologies' SOLiD platform), semiconductor-based sequencing systems (e.g., Life Technologies' Ion Torrent or Ion Proton platforms), GenapSys Gene Electronic Nano-Integrated Ultra-Sensitive (GENIUS) technology, single-molecule sequencing (e.g., Helicos systems or Pacific Biosciences systems), or nanopore-based sequencing systems (e.g., Oxford Nanopore Technologies or Roche's Genia platform (sequencing.roche.com / research---development / nanopore-sequencing.html)). Nanopores constructed using lipid bilayer membranes and protein nanopores, as well as solid-state nanopores (such as graphene-based ones), are also used. Selected single-molecule sequencing platforms allow the direct elucidation of the methylation status of DNA molecules (including N6-methyladenine, 5-methylcytosine, and 5-hydroxymethylcytosine) without bisulfite conversion (B.A. Flusberg et al. 2010 Nat Methods;7:461-465; J. Shim et al. 2013 Sci Rep:3:1389. doi:10.1038 / srep01389). Using such platforms, the methylation status of non-bisulfite-converted sample DNA (e.g., plasma or serum DNA) can be analyzed. Sequences can involve paired-end sequencing or provide single sequence reads of the entire DNA molecule.

[0063] In addition to sequencing, other techniques, such as those described above, can be used. In one embodiment, methylation profiling can be performed by methylation-specific PCR, or methylation-sensitive restriction enzyme digestion followed by PCR, or ligase chain reaction followed by PCR. In yet another embodiment, the PCR is in the form of single-molecule or digital PCR (B. Vogelstein et al. 1999 Proc Natl Acad Sci USA; 96:9236-9241). In a further embodiment, the PCR can be real-time PCR (Lo et al. Cancer Res 1999; 59(16):3899-903 and Eads et al. Nucleic Acids Res 2000; 28(8):E32). In another embodiment, the PCR can be multiplex PCR. In one embodiment, methylation profiling can be performed using microarray-based technology.

[0064] After sequencing, the sequence reads can be processed through a methylation data analysis pipeline, Methyl-Pipe (Jiang et al. PLoS One 2014;9:e100360), to map them to an artificially combined reference sequence consisting of the entire human genome (hg19), the entire EBV genome (AJ507799.2), the entire HBV genome, and the entire HPV genome. Different reference sequences can be used, and mapping can be performed for each genome rather than combining them into a single reference sequence. Sequenced reads that map to unique locations within the combined genome sequence can be used for downstream analysis. B. Targeted Bisulfite Sequencing Using Capture Probes

[0065] Certain embodiments can examine specific regions for methylation patterns in plasma EBV DNA molecules. In one embodiment, capture-enhanced targeted bisulfite sequencing can be used to analyze circulating cell-free viral DNA molecules from subjects with different EBV-related diseases or conditions. For example, capture probes can be designed to cover all or part of the CpG sites of the EBV genome. This approach can also be used for other viruses. Thus, capture probes can be designed to cover all or part of the CpG sites of the hepatitis B virus (HBV) genome, human papillomavirus (HPV) genome, and other viral / bacterial genomes. Capture probes can also be included in the same analysis to target genomic regions of the human genome.

[0066] In some embodiments, to take into account the size difference between viral genome and human genome, more probes can be designed to hybridize with viral genome sequence than the target human genome region.In another embodiment, for example, the entire viral genome can be targeted by designing an average of about 200 μM hybridization probes (for example, 200-fold tiling capture probes) that cover each viral genome region of about 200 bp in size.In one embodiment and example, for the target region in human genome, an average of 5 hybridization probes are designed to cover each region of about 200 bp in size (for example, 5-fold tiling capture probes).As an illustration, capture probes can be designed according to Figure 1.

[0067] Figure 1 shows the design of a capture probe for targeted bisulfite sequencing according to an embodiment of the present disclosure. Figure 1 shows information about the capture probe, such as the size of the capture region and the amount of tiling covered by the probe. The capture probes can be of various lengths and overlap each other. Such capture probes can be used with the SeqCap-Epi system (Nimblegen). Other embodiments may not use such capture probes.

[0068] Column 101 identifies the type of sequence, i.e., autosome for human or viral targets. Column 102 identifies the specific sequence (e.g., chromosome or specific viral genome sequence). Column 103 indicates the total length of base pairs (bp) covered by the capture probe. The capture probe may not cover the entire sequence (e.g., as shown for autosomes), but may cover the entire sequence in the case of viral genomes, for example. Column 104 indicates the capture probe depth, also called the probe fold. These numbers convey the number of probes covering any given location. For autosomes, capture probes provide 5x tiling on average. For viral targets, capture probes provide 200x tiling on average. Therefore, the number of probes for viruses will be a higher percentage / proportion per unit length than for autosomes. Such high levels of capture probe concentration for viral targets can help maximize the chances of capturing viral DNA. III. Plasma EBV DNA methylation levels in various pathological conditions

[0069] We analyzed the methylation patterns of plasma EBV DNA molecules in patients with various EBV-associated diseases / pathologies, including NPC, infectious mononucleosis, Hodgkin's lymphoma, and NK-T-cell lymphoma, as well as in apparently healthy individuals with detectable plasma EBV DNA. These apparently healthy subjects with detectable plasma EBV DNA were collected from a cohort of subjects recruited for NPC screening and divided into two groups. The first group included subjects with detectable plasma EBV DNA levels in the initial test but undetectable levels in the follow-up test, designated "transient positives." The second group included subjects with detectable plasma EBV DNA levels in both the initial and follow-up tests, designated "persistent positives."

[0070] We used targeted bisulfite sequencing with enhanced capture using specially designed capture probes. For each plasma sample analyzed, DNA was extracted from 4 mL of plasma using the QIAamp DSP DNA Blood Mini Kit. In both cases, all extracted DNA was used for sequencing library preparation using the KAPA Library Preparation Kit (Roche) or the TruSeq DNA PCR-Free Library Preparation Kit (Illumina). The adapter-ligated DNA products were subjected to two rounds of bisulfite treatment using the EpiTect Bisulfite Kit (Qiagen). 12–15 cycles of PCR amplification were performed on the bisulfite-converted samples using the KAPA HiFi HotStart Uracil+ReadyMix PCR Kit (Roche). The first PCR amplification allows for increased amounts of DNA to be captured. The input amount of DNA for the target capture reaction is suggested. The amount of input DNA from plasma (without amplification) may not be sufficient for target capture.

[0071] Next, the amplification products were captured using a SeqCap-Epi system (Nimblegen) using custom-designed probes covering the above viral and human genome regions (Figure 1). Substantial "DNA loss" can occur during the capture step. The amount of DNA after the target capture reaction may be less than the amount required for sequencing. Therefore, a second amplification step (e.g., using PCR) can amplify the amount of DNA for the subsequent sequencing step. Therefore, in some embodiments, after target capture, the capture products were enriched by 14 cycles of PCR to generate DNA libraries. The DNA libraries were sequenced on the NextSeq platform (Illumina). For each sequencing run, four to six samples with unique sample barcodes were sequenced using paired-end mode. From each DNA fragment, 75 nucleotides were sequenced from each of the two ends, although other numbers of nucleotides can also be sequenced. A. Plasma EBV DNA methylation profiles in different EBV-associated pathologies

[0072] Figure 2 shows the methylation density of CpG sites across the EBV genome in patients with infectious mononucleosis, NPC, and NK-T cell lymphoma, according to an embodiment of the present disclosure. EBV DNA methylation profiles were generated by targeted capture bisulfite sequencing of plasma EBV DNA fragments. The horizontal axis shows genomic coordinates of the EBV reference genome. The vertical axis shows methylation density at the resolution of a single CpG site.

[0073] The methylation density of CpG sites throughout the EBV genome was calculated using the formula described above.

number

[0074] Furthermore, at a relatively macroscopic level, a patient with NK-T cell lymphoma (TBR1629) showed higher heterogeneity (e.g., between genomic coordinates 50,000 and 100,000) in methylation levels across the EBV genome than two patients with NPC (TBR1392 and TBR1416). This heterogeneity appears as depressions in the methylation density plot. While NPC patients have relatively uniform density, lymphoma patients show many small valleys where density drops significantly, resulting in a comb-like structure.

[0075] DNA methylation patterns can also be analyzed at locus-specific or region-specific levels.These loci can be of any size and at least one CpG site.These loci may or may not be associated with annotated viral genes.Such region-specific methylation levels can have similar values ​​in different subjects with the same disease state, but have different values ​​in other subjects with different disease states.

[0076] In Figure 2, we define four genomic regions: region 201 (7,000–13,000), region 202 (138,000–139,000), region 203 (143,000–145,000), and region 204 (169,000–170,000). Region-specific methylation densities in regions 201 and 204 were higher in two cases of NPC (TBR1392 and TBR1416) than in a case of infectious mononucleosis (TBR1610). Conversely, region-specific methylation densities in region 203 were lower in two cases of NPC than in cases of infectious mononucleosis. Region-specific DNA methylation densities in region 203 were highest in cases of NK-T cell lymphoma (TBR1629) than in other cases of NPC and infectious mononucleosis. These results indicate that different patterns exist in the methylation profiles of plasma EBV DNA fragments on a global and locus-specific level in patients with different EBV-associated pathologies.

[0077] Thus, a low methylation level in region 201 may indicate that a subject has infectious mononucleosis. A high methylation level in region 204 may indicate that a subject has NPC. A high methylation level in region 203 may indicate that a subject has NK-T cell lymphoma. Specific threshold values ​​defining high or low (or intermediate range) can be determined for each region based on measurements of the type shown in FIG. 2. Such regions can be selected by analyzing the methylation profiles of subjects with different disease states and selecting regions with different methylation densities for the different disease states. Furthermore, measurements from multiple regions can be combined, for example, via clustering techniques or decision trees. B. Early stage NPC

[0078] Figure 3 shows the methylation profile of plasma EBV DNA in a patient (AL038) from our screening cohort with early-stage NPC (stage I) and a low concentration of plasma EBV DNA (8 copies per mL of plasma as measured by quantitative PCR). Plasma DNA was extracted from the initial blood sample. Subjects with a positive initial (baseline) test were retested 4 weeks later, which was considered a follow-up test. This patient had no symptoms of NPC at the time of blood collection, and cancer was detected by screening using real-time PCR analysis of plasma EBV DNA in a two-stage assay. Subjects with persistently positive plasma EBV DNA by real-time PCR were further confirmed using transnasal endoscopy and MRI.

[0079] Figure 3 shows that the signal is noisy; for example, some sites have 100% methylation density and some sites have very low methylation density, such as zero. To filter out such noisy behavior, embodiments can use a region methylation level measured using the combined methylation density of all sequence reads at sites within a window. For example, a 200 bp window can be used, which can reduce noise and provide smoother data. Thus, even if the concentration of EBV DNA sequences in a sample is low, the methylation level can be measured and used to distinguish between different disease states. Further data demonstrating this ability to distinguish between disease states is provided below.

[0080] In this patient, the amount of captured plasma EBV DNA fragments was relatively lower than that of two other patients (TBR1392 and TBR1416) with advanced-stage NPC and high concentrations of plasma EBV DNA. As previously mentioned, this indicates that even in cases with low EBV concentrations, methylation levels(ies) can still be used to identify specific pathologies (in this case, NPC). Furthermore, the amount of plasma EBV can be used as part of determining the level of disease (e.g., the level of cancer). C. Differences in methylation profiles between patients

[0081] The differences in methylation profiles may provide a comparison between patients with NPC and infectious mononucleosis. As previously described in Figure 2, there are different methylation patterns of plasma EBV DNA between patients with different EBV-related conditions. We analyzed the differences in methylation patterns by comparing the methylation density of CpG sites across the EBV genome between these different patients. 1. Different pathologies

[0082] Figure 4 shows the difference in methylation density of CpG sites across the EBV genome between two patients with different pathologies, according to an embodiment of the present disclosure. The horizontal axis represents the genomic coordinates of the EBV genome. The vertical axis represents the site-by-site difference in methylation between the two patients. The median methylation difference between NPC (TBR1392) and infectious mononucleosis (TBR1610) was 23.9% (IQR (interquartile range): 14.8-39.3%), indicating that NPC methylation levels across CpG sites across the EBV genome are systematically higher in NPC than in infectious mononucleosis. A similar pattern of methylation difference (median: 22.9%, IQR: 13.3-37.8%) was observed in a separate comparison between NPC (TBR1416) and infectious mononucleosis (TBR1610).

[0083] The figure above shows the difference in methylation density between an NPC patient (TBR1392) and an infectious mononucleosis patient (TBR1610). A positive value at a CpG site indicates a higher methylation density in case TBR1392 than in case TBR1610 at that particular site. A negative value indicates a lower methylation density in case TBR1392 than in case TBR1610 at that CpG site.

[0084] The figure below shows the difference in methylation density between another NPC patient (TBR1416) and the same infectious mononucleosis patient (TBR1610). This graphic representation shows an example of the analysis and comparison of plasma EBV DNA methylation patterns in different EBV-associated pathologies.

[0085] Generally, NPC patients have higher methylation, and the difference in methylation has a significant value. Such difference values ​​can be quantified in various ways, for example, by summing them to obtain an overall difference value. This overall difference value can serve as the distance between two subjects used in clustering, and each methylation value (e.g., a site index or a region level) is one data point in the multidimensional data points. 2. Same pathology

[0086] Figure 5 shows the difference in methylation density of CpG sites across the EBV genome between two patients with the same diagnosis of NPC (TBR1392 and TBR1416) according to an embodiment of the present disclosure. In general, the difference in methylation density across the EBV genome is smaller compared to previous analyses of patients with two different diseases (Figure 4). The median methylation difference between the two NPC subjects (TBR1392 vs. TBR1416) was 0.3% (IQR: -1.2-2.5%). This indicates that patients with the same diagnosis of EBV-related disease may have similar plasma EBV DNA methylation patterns. Differences in methylation density may affect several disease characteristics specific to a particular case and provide additional diagnostic or prognostic information. 3. NPC and false positives

[0087] Figure 6 shows the difference in plasma EBV DNA methylation patterns between a patient with early-stage NPC (AO050) and a subject with a false-positive plasma EBV DNA result (HB002), according to an embodiment of the present disclosure. This comparison shows the plasma EBV DNA methylation patterns of a patient with early-stage NPC and a subject without NPC who had persistently positive plasma EBV DNA results on serial testing. Both were from the screening cohort of the present invention. Plasma DNA was extracted from the initial blood sample at the time of mobilization.

[0088] As shown in Figure 6, there are differences in plasma EBV DNA methylation patterns between early-stage NPC patients (AO050) and subjects with false-positive plasma EBV DNA results (HB002). However, the number and size of the differences between NPC and IM subjects are smaller than those plotted in Figure 4. Therefore, the fact that there are differences between NPC subjects and false-positive subjects indicates the potential for improving the accuracy of cancer screening. Furthermore, the fact that the differences are on a different scale than those in IM subjects indicates any potential for distinguishing subjects with any of the three pathologies. Based on this observation, we investigated the diagnostic utility of using plasma EBV DNA methylation patterns to distinguish between the two groups (subjects with early-stage NPC and those with false-positive results), and provide the data in a later section. D. Correlation between methylation densities of similar and different patients

[0089] In addition to analyzing the differential value of methylation density at different sites, methylation densities can be plotted together to identify correlation or lack thereof.For example, each data point of a two-dimensional plot can include two methylation densities from two subjects at the same site.If methylation densities are correlated (for example, two subjects have the same pathology), the plot will show linear behavior.If methylation densities are not correlated (for example, two subjects have the same pathology), the plot will not show linear behavior.

[0090] Figures 7A-7C show the difference in plasma EBV DNA methylation profiles between two clinical cases. In Figures 7A-7C, each data point in the three graphs represents the methylation density of a CpG site in the entire EBV genome of one patient (on the x-axis) and the methylation density of the same CpG site in the corresponding other patient (on the y-axis).

[0091] Figures 7A and 7B show methylation densities between two patients with different diseases (one NPC and one infectious mononucleosis). As can be seen, the methylation densities are not correlated. The NPC subjects consistently have high methylation densities (e.g., 80% or higher) that are not consistent with the IM subjects, resulting in a horizontal band at the top. This behavior indicates that the two subjects are different, e.g., have different disease states. Different disease states can include those with and without disease.

[0092] Figure 7C shows the methylation density between two different patients with NPC. In Figure 7C, a diagonal trend line (with a slope approximately equal to 1) can be observed, suggesting that the methylation density of each CpG site is similar between two different patients with NPC. This graphical pattern is not observed in Figures 7A and 7B. These results again suggest that patients with different EBV-related diseases have different methylation profiles of plasma EBV DNA fragments. Embodiments can use such differential methylation characteristics of sites (or regions of sites) to identify sites / regions with different methylation densities between different disease states and use those sites / regions to determine methylation level(s) for distinguishing disease states. IV. Identification of EBV-associated pathologies using plasma EBV DNA methylation patterns

[0093] For a systematic comparison of plasma EBV DNA methylation profiles in different EBV-associated pathologies, we used the "methylation percentage" (an example of a kind of "methylation density") for each case. The methylation percentage of plasma EBV DNA fragments can be derived using the following equation:

number

[0094] The methylation percentage can be determined as an example of a single methylation level across the entire EBV genome. The aggregate number of EBV DNA molecules methylated at a specific set of CpG sites can be used to determine the methylation percentage as the genome-wide methylation level, along with normalization by a measurement that includes other DNA molecules, such as a measure of volume, intensity corresponding to other DNA molecules, or the number of other DNA molecules. In one embodiment, the inventors calculated the methylation percentage of plasma EBV DNA molecules based on the CpG sites in the EBV genome covered by the capture probes of the present invention.

[0095] Figure 8 shows the methylation percentage of plasma EBV DNA based on all CpG sites covered in subjects with infectious mononucleosis (IM) (n=2 patients), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6), according to an embodiment of the present disclosure. Figure 8 shows boxplots of five different pathologies. As shown, the medians are sufficiently separated to distinguish between different pathologies. For example, a reference level of approximately 79% can distinguish between patients who are persistently EBV positive and those with NPC.

[0096] Overall, analysis of the genome-wide aggregate methylation percentages of plasma EBV DNA was able to distinguish between different EBV-associated diseases / conditions (p-value = 8.52e-05, one-way ANOVA test). Four of the six NPC patients from the screening cohort had early-stage NPC (stage I or II). Therefore, even early-stage disease states can be distinguished from subjects without the condition (e.g., persistently positive subjects). Different reference levels can be used to distinguish between different conditions; for example, less than about 57% can be used to identify IM, and 57%-63% can be used to identify transiently positive subjects. In some embodiments, for a given methylation level, two or more conditions from a broader set of conditions can be identified (e.g., lymphoma and transiently positive cases can be identified in the 57%-63% range). In such situations, a probability can be assigned to each condition. For example, the methylation of a test subject can be compared to two groups of reference subjects, one group suffering from condition A (e.g., lymphoma) and another group suffering from condition B (e.g., infectious mononucleosis). The number of standard deviations from the mean of the two groups of reference subjects can be determined. Based on the number of standard deviations, probabilities of having the two conditions can be calculated. These probabilities can be used to determine the relative likelihood of having the two conditions.

[0097] The particular reference value(s) selected may depend on the particular manner in which the methylation level is determined. For example, the methylation percentage is the number M of methylated DNA molecules (e.g., determined using the number of reads or intensity) divided by the number U of unmethylated DNA molecules (e.g., determined using the number of reads or intensity), e.g., M / U. Other scaling or summation factors may modify the particular reference value. The selected reference level is applicable as long as the methylation level of the present sample is determined in the same manner as the methylation level of the reference sample. The reference level may also depend on the site selected to measure the methylation level, as shown in the results below.

[0098] While the ability to distinguish between different pathologies using methylation levels may depend on the number of DNA molecules detected at a site, the results presented herein indicate that the number of EBV DNA molecules can be relatively small. For example, in Figure 8, the 5% value of plasma EBV DNA molecules in different subjects is 44, with a minimum of 26. In various embodiments, the number of cell-free viral DNA molecules can include at least 10 cell-free DNA molecules, e.g., 20, 30, 40, 50, 100, or 500 cell-free DNA molecules. In other embodiments, the total number of cell-free DNA molecules is analyzed for a subject, and a particular viral genome can be at least 1,000 cell-free DNA molecules or more (e.g., at least 10,000, at least 100,000, or at least 1,000,000). B. Differentially methylated regions (DMRs)

[0099] Instead of using all sites covered by the capture probe, only specific sites can be used. These sites can be determined by analyzing all sites and selecting sites with specific characteristics, such as methylation levels that differ between specific pathologies. Sites can be analyzed individually or collectively by region, for example, two or more sites can be assigned to one region and the methylation level determined for that region. Sites and regions can span the viral genome, for example, not be limited to one site / region, but can be the entire genome. For example, at least one site / region can be used every 1 kb, 2 kb, 5 kb, or 10 kb.

[0100] Therefore, some embodiments can calculate the methylation percentage based on CpG sites within differentially methylated regions (DMRs). It has previously been shown that the methylation patterns of plasma EBV DNA differ in different EBV-related diseases. Therefore, DMRs should be present, and within the DMRs, methylation levels differ between different diseases / pathologies. In addition to using individual sites, non-overlapping windows of different sizes can be used. For example, the size of the non-overlapping region can be set to, but is not limited to, 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 800 bp, and 1000 bp. In another example, each CpG site can be analyzed individually, for example, without using binding site regions.

[0101] DMRs can be selected to have specific methylation levels for subjects with different disease states. For example, a DMR can be defined when the methylation rate of CpG sites in the region is less than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% in one or more cases of a disease / pathology and more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% in one case of another disease / pathology. In yet another embodiment, two or more cases per disease can be used to define a DMR. Also, cutoff criteria for more than two diseases / pathologies can be used, such as different ranges of methylation levels for each condition. 1.DMR with IM<50% and NPC>80%

[0102] To demonstrate mining such DMRs, we randomly selected one patient with infectious mononucleosis (TBR1610) and one patient with NPC (TBR1392). Here, we first set non-overlapping windows (bins of positions 1–500, 501–1000, etc.) of 500 base pairs (bp) across the entire EBV genome. Within the 500-bp region, we calculated the average methylation percentage of all CpG sites within the region in patients with IM (TBR1610) and NPC (TBR1392). In this example, a 500-bp region may meet the first selection criterion for DMRs if the average methylation percentage of all CpG sites within the region is less than 50% in the case of infectious mononucleosis (TBR1610) and greater than 80% in the case of NPC (TBR1392).

[0103] Figure 9 illustrates mining of differentially methylated regions (DMRs) that meet the first selection criteria according to an embodiment of the present disclosure. Figure 9 corresponds to the two cases shown in Figure 7A. Each data point in Figure 9 represents the methylation percentage of a 500-bp region across the EBV genome in an IM subject (x-axis) and the corresponding methylation percentage of the same 500-bp region in an NPC subject (y-axis). Using this first selection criterion defined above, we identified a total of 39 DMRs consisting of 821 CpG sites (approximately 10% of all CpG sites in the EBV genome captured by the probes of the present invention). These 39 DMRs are located in the upper left corner of Figure 9.

[0104] In Figure 9, the methylation fraction of NPC subjects is shown on the vertical axis, and the methylation fraction of IM subjects is shown on the horizontal axis. Vertical line 901 indicates the 50% cutoff for methylation fraction of IM cases. Horizontal line 902 indicates the 80% cutoff for NPC cases. Thus, the region in the upper left section (generally marked as 910) corresponds to the DMR in this example.

[0105] Figure 10 is a table listing the genomic coordinates of 39 differentially methylated regions that meet the criteria set forth in Figure 9. Column 1001 lists the viral genome, in this example EBV. Column 1002 shows the start genomic coordinate of the reference EBV genome. Column 1003 shows the end genomic coordinate of the reference EBV genome. Methylation Density Column 1004: Methylation density in IM and NPC subjects.

[0106] These DMRs can be used to determine the methylation level(s) of other subjects. In one embodiment, the methylation state of each sequence read covering one site in the DMR set is used to determine the proportion corresponding to the DMR set. This methylation proportion can be determined in a manner similar to that shown in FIG. 8, but using a subset of sequence reads, i.e., those corresponding to sites within the set of DMRs. In other embodiments, individual methylation levels can be determined for each DMR. The methylation levels of subjects can form multidimensional data points, and for example, clustering techniques or reference planes (hyperplanes in multidimensional space) can separate subjects with different disease states or different classifications / levels of disease states. Other analytical methods can be used to distinguish these disease states, including, but not limited to, naive Bayes, random forests, decision trees, support vector machines, k-nearest neighbors, K-means clustering, Gaussian mixture models (GMMs), density-based spatial clustering, hierarchical clustering, logistic regression classification, and other supervised and unsupervised classification or regression methods.

[0107] Figure 11 shows plasma EBV DNA methylation percentages based on 821 CpG sites within the 39 DMRs described above (Figure 8) in the same group of subjects with infectious mononucleosis (IM) (n=2), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6), according to an embodiment of the present disclosure. These are the same subjects used in Figure 8. The methylation percentages are determined as a single value using sequence reads that correspond to (e.g., align with) the 821 sites within the set of 39 DMRs.

[0108] Figure 11 shows boxplots of five different pathologies. As shown, the median values ​​can distinguish between different pathologies. For example, a reference level of approximately 75% can distinguish between patients with persistently positive EBV and those with NPC. A statistically significant difference in methylation rates between the different groups can be observed (p-value = 1.83e-05, one-way ANOVA test), which is superior to Figure 8.

[0109] Figure 11 has some differences from Figure 8. For example, the spread between IM and NPC values ​​is smaller because the DMRs were specifically selected to be in a specific range for these subjects. These results indicate that different range criteria for methylation levels indicating different pathologies can provide a more selective technique for distinguishing different classifications of subjects. Furthermore, the difference between NPC subjects and persistently positive subjects is larger than in Figure 8. 2.DMR with IM<80% and NPC>90%

[0110] Similar to the previous section, we performed differential methylation analysis based on differentially methylated regions to distinguish early-stage NPC patients from non-NPC patients with detectable plasma EBV DNA. All analyzed NPC patients and non-NPC subjects were identified through prospective screening cohorts (Chan et al. N Engl J Med 2017;377:513-522). Two NPC patients (TBR1416 and FD089) and one infectious mononucleosis patient (TBR1748) were randomly selected for DMR mining. In yet another embodiment, other EBV-related diseases or conditions, including non-EBV subjects with detectable plasma EBV DNA, may be used for DMR mining. Non-overlapping windows of 500 base pairs (bp) in size were set across the EBV genome.

[0111] Figure 12 illustrates the mining of differentially methylated regions (DMRs) that meet the second selection criteria, according to an embodiment of the present disclosure. In contrast to the previous section, mining of DMRs included two or more cases per disease (NPC in this analysis). The second selection criteria correspond to (1) non-overlapping, contiguous windows 500 bp in size, and (2) a methylation percentage of CpG sites within the selected region that is less than 80% in the selected infectious mononucleosis case (TBR1748) and greater than 90% in both nasopharyngeal carcinoma cases (TBR1416 and FD089). Each data point in Figure 12 represents the average methylation percentage of all CpG sites within a 500-bp region across the EBV genome in an IM subject (x-axis) and the corresponding average methylation percentage of the same region in two NPC subjects (y-axis).

[0112] In Figure 12, the methylation percentage of NPC subjects is shown on the vertical axis, and the methylation percentage of IM subjects is shown on the horizontal axis. Vertical line 1201 indicates the 80% cutoff for the methylation percentage of IM cases. Horizontal line 902 indicates the 90% cutoff for NPC cases. Using this second selection criterion defined above, a total of 46 DMRs consisting of 1,520 CpG sites (approximately 20% of all CpG sites in the EBV genome captured by the probes of the present invention) were identified. The 46 DMRs are shown in the upper left section (generally marked as region 1210).

[0113] 13 shows plasma EBV DNA methylation percentages based on the 46 DMRs specified in FIG. 12 in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to an embodiment of the present disclosure. Each data point corresponds to a different subject. The classification of transiently positive and persistently positive is determined as described above, i.e., based on the number of EBV DNA reads collected twice from the sample.

[0114] Targeted bisulfite sequencing was used to analyze 117 non-NPC subjects with transiently positive plasma EBV DNA, 39 non-NPC subjects with persistently positive plasma EBV DNA, and 30 NPC patients. All non-NPC and NPC subjects were recruited from prospective screening cohorts. Plasma EBV DNA methylation rates were compared among the three groups based on the DMRs defined above (Figure 12). Methylation rates were determined as a single value using sequence reads corresponding to (e.g., aligned with) 1,520 sites within a set of 46 DMRs.

[0115] In one embodiment, different weights can be assigned to different DMRs for the calculation of the aggregate methylation percentage. Such weighting can be implemented in various ways, for example, by a scaling factor for each methylated sequence read at a site (e.g., multiplying by a factor greater than 1 to give a region a higher weight), or by applying a scaling factor to the methylation percentage of a region, thereby providing a weighted average of the methylation percentages of the region. In this example, equal weights are applied to all defined DMRs.

[0116] The mean methylation percentage of plasma EBV DNA based on 46 DMRs in the NPC group (mean = 88.3%) was significantly higher than the mean methylation percentages of the other two non-NPC groups with transiently positive (mean = 65.3%) and persistently positive (mean = 71.1%) plasma EBV DNA (p < 0.0001, Kruskal-Wallis test). Therefore, NPC patients can be distinguished from non-NPC subjects with detectable plasma EBV DNA (transiently positive or persistently positive) based on the difference in plasma EBV DNA methylation profile (e.g., expressed as DMR-based plasma EBV DNA methylation percentage).

[0117] In this example and other embodiments described herein, the reference level for distinguishing between classifications can be determined in various ways. In one embodiment, the cutoff value (reference level) used to distinguish between NPC and non-NPC subjects can be the lowest EBV DNA methylation percentage among the NPC patients (training set) under analysis. In other embodiments, the cutoff value can be determined, for example, as the mean EBV DNA methylation percentage of NPC patients minus one standard deviation (SD), the mean minus two SDs, or the mean minus three SDs. In yet other embodiments, the cutoff can be determined using a receiver operating characteristic (ROC) curve by nonparametric methods, including, but not limited to, 100%, 95%, 90%, 85%, and 80% of the NPC patients under analysis.

[0118] In this example, a cutoff value of 80% for the methylation percentage could be set to achieve a sensitivity of over 95% for NPC detection. Using this 80% cutoff value, 29 of 30 NPC patients, 16 of 119 non-NPC subjects with transiently positive plasma EBV DNA, and 6 of 39 non-NPC subjects with persistently positive plasma EBV DNA had plasma EBV DNA methylation percentages determined within the 46 DMRs defined by targeted bisulfite sequencing that were higher than the cutoff value (80%). The calculated sensitivity, specificity, and positive predictive value were 96.7%, 85.9%, and 58.5%, respectively. C. Representative methylation consensus regions for the same pathology

[0119] In the following example, we aimed to identify regions with similar methylation densities between subjects with the same disease state. Such regions are defined as representative methylation consensus regions of the disease state. In some embodiments, such criteria can also be combined with criteria for differential methylation between disease states.

[0120] To demonstrate the identification of "representative" methylation consensus regions, two NPC patients (TBR1392 and TBR1416) were randomly selected. Here, the EBV genome was divided into 500-bp non-overlapping regions. In other embodiments, overlapping regions of different sizes may be set. For example, the size of the overlapping region may be set to 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 600 bp, 800 bp, and 1000 bp. Within the 500-bp region, the average methylation percentage of all CpG sites within the region was calculated for the two patients with NPC (TBR1392 and TBR1416).

[0121] 14 illustrates mining of representative methylation consensus regions that meet the third selection criteria, according to an embodiment of the present disclosure, corresponding to (1) non-overlapping, contiguous windows of 500 bp in size, (2) a difference in overall methylation density of less than 1% of the region between the two NPC cases, and (3) an overall methylation density of more than 80% of the region in both cases.

[0122] In other embodiments, a "representative" methylation consensus region may be defined if the difference in methylation percentage between two NPC patients is less than 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% between the two NPC cases, and the methylation percentage is greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, or 90%. More than two subjects per disease may be used to define a "representative" methylation consensus region, for example, the methylation percentage of each subject is within a certain similarity cutoff (e.g., 1%) and within a certain range of methylation percentages (e.g., greater than 80%).

[0123] Each data point in Figure 14 represents the methylation percentage of a 500 bp region across the EBV genome in an NPC subject (on the x-axis) and the methylation percentage of the same region in other NPC subjects (on the y-axis). Using the selection criteria defined above, we identified 79 regions. These 79 DMRs can be found in region 1410 in Figure 14.

[0124] In Figure 14, the methylation percentage of a first NPC subject is shown on the vertical axis, and the methylation percentage of a second NPC subject is shown on the horizontal axis. Vertical line 1401 indicates the 80% cutoff for the methylation percentage of IM cases. Horizontal line 1402 indicates the 80% cutoff for NPC cases. Thus, the region in the upper right section (generally marked as region 1410) corresponds to the DMR in this example.

[0125] Figure 15 shows the methylation density of plasma EBV DNA based on the "representative" methylation consensus region described above (Figure 12) in the same group of subjects with infectious mononucleosis (IM) (n=2), EBV-associated lymphoma (n=3), transiently positive plasma EBV DNA (n=3), persistently positive plasma EBV DNA (n=3), and NPC (n=6), according to an embodiment of the present disclosure. A statistically significant difference in the methylation rate between the groups was observed (p-value=0.00371, one-way ANOVA test). Based on these data, the status of the test sample can be determined by analyzing the methylation density of plasma DNA at the representative methylation consensus region. For example, a methylation density of 90% indicates that the sample is from an NPC patient, and a methylation density of 80% indicates that the sample is from an EBV-positive lymphoma patient. D. Single CpG site analysis

[0126] In addition to calculating aggregate methylation levels based on a set of sites (e.g., as described in the section above), embodiments can calculate methylation percentages based on individual CpG sites within the EBV genome. To identify individual CpG sites with methylation levels specific to different EBV-associated diseases, we pooled sequence data from plasma DNA reads from three subjects with persistently positive EBV DNA but without NPC, increasing the sequencing depth by sevenfold. Next, we compared the methylation percentages of all CpG sites between the pooled sequence data of three subjects with persistently positive EBV DNA and three NPC patients (AO050, TBR1392, and TBR1416). Pooling the sequence data means that the methylation percentages were determined as if all sequence reads were from the same subject.

[0127] In various embodiments, individual CpG sites may be defined as having differential methylation levels if the methylation rate of the CpG site is less than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% in one case (subject) of a disease and greater than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% in one case of another disease. Criteria may be applied to more than one case per disease to define DMRs, e.g., as described above for the example using regions, to allow for different extents of different disease states, as well as specific separation between subjects with different (greater) disease states or the same disease state (within a threshold).

[0128] Appendix A shows a list of individual CpG sites across the EBV genome with differential methylation levels, where the difference in methylation percentage at these CpG sites between the pooled sequence data of three subjects with persistently positive EBV DNA and three NPC patients is greater than 20%. Sites marked with * have a difference of greater than 40%, ** greater than 60%, and *** greater than 80%. In other embodiments, CpG sites with differential methylation levels may be defined by a difference in methylation percentage of greater than 20%, 30%, 50%, 70%, or 90%. Here, several exemplary sites with a difference of greater than 60% are analyzed.

[0129] Figure 16 shows examples of CpG sites with methylation percentages across sites greater than 80% in pooled sequencing data from three cases with persistently positive plasma EBV DNA and a mean value of less than 20% across three subjects with NPC, according to embodiments of the present disclosure. Methylation percentages across these sites in two patients with infectious mononucleosis were also included. Methylation percentages are shown on the vertical axis, and the horizontal axis lists eight individual sites that meet the criteria. Sites are numbered consecutively.

[0130] As shown in the figure, all sites provide good separation between non-NPC subjects (including persistently positive and IM) and two of the NPC subjects (TBR1416 and TBR1392). For NPC subject AO050, sites 1-4, 7, and 8 provide good separation, but sites 5 and 6 do not. Therefore, individual differentially methylated sites (e.g., not simply regions) determined to distinguish between two pathologies can also be used to distinguish between one pathology (NPC) and two or more pathologies.

[0131] Figure 17 shows examples of CpG sites with methylation percentages across less than 20% of sites in pooled sequencing data from three cases with persistently positive plasma EBV DNA and greater than 80% of sites in three subjects with NPC, according to embodiments of the present disclosure. Also shown are methylation percentages across these sites in two patients with infectious mononucleosis. This criteria is the opposite of the criteria used in Figure 16. Methylation percentages are shown on the vertical axis, and the horizontal axis lists 22 individual sites that meet the criteria. Sites are numbered consecutively.

[0132] As shown, all sites provide good separation between non-NPC subjects (including persistently positive and IM) and NPC subjects. This demonstrates that individual sites can be used to distinguish between classifications. Additionally, sites can be selected across the entire viral genome, as opposed to within the same contiguous region of a specific length. For example, selecting more sites allows for greater statistical precision, allowing for detection of more EBV DNA fragments.

[0133] In one embodiment, multiple CpG sites with differential methylation levels (e.g., those specified in Appendix A and Figures 16 and 17) are selected to be within a specified distance from each other. In this way, individual viral DNA fragments can each cover multiple sites. For example, the specified distance can be 150 bp, which is approximately the typical size of a plasma DNA molecule. In such a situation, targeted amplification of specific regions with multiple differentially methylated CpG sites by PCR amplification would be possible. This targeted analysis would be less costly than using a genome-wide analysis approach.

[0134] Therefore, in various embodiments, the analysis of methylation patterns can be based on the genomic regions of the entire viral genome or on individual CpG sites. In such regional analysis, the viral genome can be divided into different regions based on genomic coordinates, and each region contains all the CpG sites within that region. Alternatively, differentially methylated CpG sites can be first selected and then merged within the region to form DMRs. In another example, all CpG sites can be included in the calculation of the methylation density of a region without preselecting informative sites. E. Hierarchical Clustering Analysis

[0135] Some embodiments may use methylation levels to distinguish between classifications. In one example, clustering techniques may be used. In such clustering techniques, multiple methylation levels can be determined for each subject, for example, the methylation levels of different regions (each containing one or more sites), the methylation levels of individual sites, or a combination thereof. In some embodiments, clustering may be hierarchical.

[0136] The set of methylation levels can form a vector representing multidimensional data points with a length equal to the number of methylation levels. As described herein, the regions (bins) can be of various sizes. Clustering analysis can also be based on non-overlapping contiguous bins of different sizes. For example, the bin size can be defined as 50 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1000 bp. Therefore, clustering analysis can be performed based on the comparison of the methylation levels of different sites / bins with the corresponding methylation levels between different subjects.

[0137] Figure 18 shows a cluster dendrogram using hierarchical clustering analysis based on methylation pattern analysis of plasma EBV DNA for six NPC patients (including four patients with early-stage disease from the screening cohort of the present invention), two patients with extranodal NK-T cell lymphoma, and two patients with infectious mononucleosis, according to an embodiment of the present disclosure. In this example, the hierarchical clustering analysis was based on a comparison of the methylation percentages of CpG sites within non-overlapping contiguous regions 500 bp in size. Clustering can group different subjects based on differences in the methylation percentages of different regions, and the differences can be combined to provide a distance.

[0138] In Figure 18, distance is shown along the top horizontal axis. For example, if the methylation densities of different loci correspond to vectors representing multidimensional points (the distance is between two multidimensional points), the distance can be determined as the sum of the differences in each methylation density (percentage). Two subjects can be paired at a point equal to the distance between them. When a new patient is tested, the new multidimensional point (or other level) of methylation percentage is used to determine the closest one or closest subgroup of reference subjects (e.g., those shown in Figure 18), as shown in a node (e.g., node 1801 or 1802). Identifying the closest reference subject or reference node can provide classification.

[0139] Cluster dendrogram 1800 shows that NPC subjects clustered together incrementally (all NPC subjects are clustered in a subgroup at node 1803) and did not include subjects with other pathologies. This demonstrates the ability to distinguish NPC subjects from IM and lymphoma subjects. Similarly, IM subjects are grouped together. Notably, two patients with NK-T cell lymphoma did not cluster together. Patient 1629 had stage IV disease, and patient 1713 had stage I disease. This may indicate that methylation patterns can progress across different stages of the same disease. One potential application of this would be using methylation profiles for patient staging and prognosis.

[0140] The feasibility of distinguishing between different EBV-associated diseases was investigated based on clustering analysis of plasma EBV DNA methylation patterns. For example, after specific patterns are identified, embodiments can include or exclude diseases. Other classification algorithms can also be used, including, but not limited to, principal component analysis, linear discriminant analysis, logistic regression, machine learning models, k-means clustering, k-nearest neighbors, and random decision forests.

[0141] Figure 19 shows a cluster dendrogram using hierarchical clustering analysis based on methylation pattern analysis of plasma EBV DNA for six NPC patients (including four early-stage NPC patients from the screening cohort of the present invention) and three non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. Plasma DNA was extracted from the initial blood sample at the time of mobilization. In this example, the hierarchical clustering analysis was based on comparison of the methylation percentages of CpG sites within non-overlapping contiguous bins of 500 bp in size.

[0142] In Figure 19, distance is also shown along the top horizontal axis. Distance can be determined as described above. As in Figure 18, two subjects can be paired at a point equal to the distance between them. As shown, NPC subjects are clustered with other NPC subjects, for example, at nodes 1901 and 1902. Cluster dendrogram 1900 shows that NPC subjects are clustered together incrementally (all NPC subjects are clustered in a subgroup at node 1903) and do not include subjects with other pathologies (i.e., persistent positive in this example). Similarly, persistent positive subjects are grouped first. This demonstrates the ability to distinguish NPC subjects from persistent positive subjects.

[0143] Thus, we demonstrated the feasibility of distinguishing early-stage NPC patients from non-NPC subjects with false-positive plasma EBV DNA based on methylation pattern analysis of plasma EBV DNA from the first blood sample, without the need for serial analysis. This means that accurate classification can be achieved with a single measurement, rather than requiring multiple measurements at different times. This potentially saves medical costs by arranging serial blood testing and further investigations for confirmation.

[0144] Figure 20 shows a heatmap 2000 showing the methylation levels of all non-overlapping 500-bp regions in the entire EBV genome in patients with nasopharyngeal carcinoma, NK-T cell lymphoma, and infectious mononucleosis. The methylation level of each window was calculated as the average methylation density of all CpG sites (without selection) within the window. Plasma EBV DNA methylation patterns were analyzed from five patients with NPC, three patients with NK-T cell lymphoma, and three patients with infectious mononucleosis. In this example, the methylation patterns are analyzed by the methylation percentage of all non-overlapping 500-bp windows. The color key / histogram 2050 indicates different methylation percentages with different colors: white indicates near zero, yellow (light gray) indicates low (e.g., approximately 20%), orange (medium gray) indicates moderate (e.g., approximately 50%), and red (dark gray) indicates high (e.g., approximately 80%). Color key / histogram 2050 also shows a histogram of the number of regions with a particular methylation level.

[0145] The heatmap 2000 shows the methylation levels across all non-overlapping 500-bp windows on the EBV genome for all cases. Each row represents one 500-bp region, and the color represents its methylation level. Each column represents one case. The cluster dendrogram 2010 shows the clustering of different cases. Different cases with the same diagnosis clustered together. For example, all NPC cases are clustered on the right and show high methylation levels, as indicated by the dark red (dark gray). Lymphoma subjects cluster in the center and show a mixture of yellow (light gray) (low methylation levels) and red (dark gray) (high methylation levels). The cluster of infectious mononucleosis samples is shown on the left and shows relatively low methylation levels, as indicated by the light yellow (light gray).

[0146] Figure 20 shows the feasibility of predicting EBV-associated disease through analysis of methylation patterns of plasma EBV DNA. Furthermore, higher accuracy is seen when using DMRs as opposed to all regions across the EBV genome.

[0147] In other embodiments, the methylation pattern can be derived, for example, by site-by-site, through the methylation percentage across all CpG sites in the entire EBV genome (as a genome-wide analysis). In another embodiment, for the prediction of EBV-associated disease or pathology, different weights can be assigned to individual CpG site(s) and / or DMR(s). Such weighting can be implemented in various ways, as described above, for example, by applying a scaling factor (weight) to the intermediate methylation level to obtain a weighted average value for the region or the entire genome. Here, all CpG sites under analysis were assigned equal weights. V. Use of Numbers and Sizes

[0148] In addition to using the methylation level of cell-free viral DNA fragments in a cell-free sample to distinguish between subjects with different disease states and / or levels of disease, some embodiments may use the size of the cell-free viral DNA fragments in the cell-free sample. Some embodiments may also use the number (e.g., percentage) of cell-free viral DNA fragments in the cell-free sample. Various embodiments may use a combination of different techniques, for example, by requiring the same classification of subjects using each technique. For example, any combination of a) the percentage of plasma DNA fragments that align to EBV, b) the size profile of plasma EBV DNA fragments, and c) the methylation profile of plasma EBV DNA may be used for classification. When combining different techniques to achieve the desired diagnostic sensitivity and specificity, different thresholds may be employed for classification. A. Size profile analysis of plasma EBV DNA fragments

[0149] In addition to demonstrating the feasibility of methylation analysis of plasma EBV DNA fragments, the size of each plasma EBV DNA fragment was estimated based on the coordinates of the outermost nucleotides at both ends of the EBV genome. The size distribution of cell-free EBV DNA fragments varies with different pathologies (i.e., different patterns), making it possible to distinguish subjects with different pathologies, i.e., different pathology levels. These different size patterns can be quantified using various size metrics, such as the ratio of the amount of viral DNA of one size (e.g., a first size range) to the amount of viral DNA of another size (e.g., a second size range). For example, size ratios can be used to compare the proportion of plasma EBV DNA reads within a specific size range (e.g., 80–110 base pairs) normalized to the amount of autosomal DNA fragments within the same size range between subjects with different pathologies.

[0150] Figure 21 shows size profiles of the size distribution of sequenced plasma DNA fragments mapped to the EBV genome and human genome in two NPC patients (TBR1392 and TBR1416) and two infectious mononucleosis patients (TBR1610 and TBR1661), as well as three non-NPC subjects (AF091, HB002, and HF020) with persistently positive plasma EBV DNA in serial analyses, according to an embodiment of the present disclosure. The plot shows size (bp) along the horizontal axis and the frequency (percentage) of DNA at the specified size along the vertical axis. The size distribution of human genomic DNA is shown in blue (gray), e.g., size distribution 2104. The EBV size distribution is shown in red (black), e.g., size distribution 2103.

[0151] We observed differences in the size profile patterns of plasma EBV DNA fragments aligned to the EBV genome and those aligned to the autosomal genome. For example, NPC subjects exhibit a shift to smaller cell-free EBV DNA fragments with a peak of approximately 160 bp and a similar amount of fragments as human DNA at the lower end, whereas IM subjects have a greater amount of EBV DNA than human DNA at less than 100 bp. Persistently positive subjects not only exhibit a significant fluctuation up and down as the size increases, but also a more pronounced peak shift compared to NPC subjects. These differences can be used to distinguish subjects with different pathologies, for example, to distinguish subjects with NPC from subjects with false-positive plasma EBV DNA results.

[0152] Comparing the proportion of plasma EBV DNA reads within a particular size range (e.g., 80-110 bp) between individuals can normalize the amount of plasma EBV DNA fragments to the amount of autosomal DNA fragments within the same size range. This metric is an example of a size ratio. The size ratio can be defined as the proportion of plasma EBV DNA within a particular size range divided by the proportion of a reference set of sequences (e.g., DNA fragments from human autosomes) within the corresponding size range. Various size ratios may be used. For example, the size ratio for fragments between 80 and 110 base pairs is as follows:

number

[0153] Figure 22 shows the size ratios in six NPC patients and three subjects persistently positive for plasma EBV DNA, according to an embodiment of the present disclosure. A statistically significant difference could be observed between the size ratios of the two groups of subjects (p-value = 0.02, Mann-Whitney test). An example of a reference size value for distinguishing between NPC subjects and persistently positive subjects could be 2 to 4 (e.g., 3) for this particular size ratio.

[0154] Figure 23 shows EBV DNA size ratios in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to an embodiment of the present disclosure. The mean EBV DNA size ratio (80-110 bp) of the NPC group (mean = 1.9) was significantly lower than the median size ratios of the other two non-NPC groups, transiently positive (mean = 4.3) and persistently positive (mean = 4.8) plasma EBV DNA (p<0.0001, Kruskal-Wallis test).

[0155] Therefore, NPC patients could be distinguished from non-NPC subjects with detectable plasma EBV DNA (transiently positive or persistently positive) based on differences in the size profile of plasma EBV DNA, expressed, for example, as the EBV DNA size ratio. In this example, a cutoff value of 3 was used to achieve 90% detection sensitivity. Using a cutoff value of 3, 27 of 30 NPC patients, 23 of 117 non-NPC subjects with transiently positive plasma EBV DNA, and 7 of 39 non-NPC subjects with persistently positive plasma EBV DNA passed the cutoff, and their plasma EBV DNA size ratios were lower than the cutoff value. The calculated sensitivity, specificity, and positive predictive value were 90%, 80.8%, and 49.2%, respectively.

[0156] The selection of the cutoff (reference) value can be determined in various ways. In one embodiment, the cutoff value of the EBV DNA size ratio can be selected as any value above the highest EBV DNA size ratio of the NPC patients under analysis (e.g., in the training set). In other embodiments, the cutoff value can be determined, for example, as the mean EBV DNA size ratio of the NPC patients plus one standard deviation (SD), the mean plus two SDs, or the mean plus three SDs. In yet other embodiments, the cutoff can be determined using a receiver operating characteristic (ROC) curve by a nonparametric method, for example, including 100%, 95%, 90%, 85%, and 80% of the NPC patients under analysis.

[0157] Other definitions of size ratios or other statistical values ​​of size distributions will yield different values ​​for each subject and therefore have different reference values ​​for distinguishing subjects. For example, different size ranges can be used, or a subset of chromosomes can be used for autosomal DNA fragments, or no autosomal DNA can be used at all. Various statistical values ​​of the size distribution of nucleic acid fragments can be determined. For example, the mean, mode, median, or average of the size distribution can be used. Other statistical values, such as the cumulative frequency of a given size, or various ratios of the amount of nucleic acid fragments of different sizes, can be used. The cumulative frequency can correspond to the proportion (e.g., percentage) of DNA fragments that are less than or greater than a given size. Therefore, any normalization factor in the denominator (if used) can be relative to the amount of EBV DNA in different size ranges. The statistical values ​​provide information about the distribution of DNA fragment sizes for comparison with one or more size thresholds for healthy control subjects or other pathological conditions. Those skilled in the art will know how to determine such thresholds based on the present disclosure. Other examples of size ratios can be found in U.S. Patent Publication Nos. 2011 / 0276277, 2013 / 0237431, and 2016 / 0217251. B. number

[0158] In addition to analyzing the methylation percentage of plasma EBV DNA, we also analyzed the percentage of plasma EBV DNA reads from targeted bisulfite sequencing. The percentage of cell-free EBV DNA reads can be determined in various ways, such as as a percentage of all DNA reads from the human genome and several viral genomes, or from only the human genome and viral genome under analysis. In the former example, the combined reference sequence can include the entire human genome (hg19), the entire EBV genome (AJ507799.2), the entire HBV genome, and the entire HPV genome. In various examples, the percentage can be determined based on the number of reads that align to the viral genome under analysis compared to all other DNA reads or only those that align. Using reference genomes for the human genome and several viral genomes (e.g., uniquely or with a specific number of mismatches), we compared the percentage of plasma EBV DNA reads between three groups of NPC patients and non-NPC subjects with transiently positive and persistently positive plasma EBV DNA.

[0159] Figure 24 shows the percentage of plasma EBV DNA reads (plasma DNA reads mapped to the EBV genome) among all sequenced plasma DNA reads in non-NPC subjects with transiently positive plasma EBV DNA, non-NPC subjects with persistently positive plasma EBV DNA, and NPC patients, according to an embodiment of the present disclosure. The mean percentage of plasma EBV DNA reads in the NPC group (mean = 0.075%) was significantly higher than the mean percentages of transiently positive (mean = 0.003%) and persistently positive (mean = 0.052%) plasma EBV DNA in the other two non-NPC groups (p < 0.0001, Kruskal-Wallis test). Therefore, NPC patients can be distinguished from non-NPC subjects with detectable plasma EBV DNA (transiently positive or persistently positive) based on differences in the amount of plasma EBV DNA (i.e., percentage of plasma EBV DNA).

[0160] Cutoff value 4.5x10 -6In this example, the plasma EBV DNA percentage by targeted bisulfite sequencing was higher than this cutoff value in all 30 NPC patients, 79 of 119 non-NPC subjects with transiently positive plasma EBV DNA, and 34 of 39 non-NPC subjects with persistently positive plasma EBV DNA. The calculated sensitivity, specificity, and positive predictive value were 100%, 27.6%, and 22.1%, respectively.

[0161] The selection of the cutoff (reference) value can be determined in various ways. In one embodiment, the cutoff value for the plasma EBV DNA percentage can be selected as any value below the lowest percentage of NPC patients under analysis. The cutoff can be set to capture all NPC patients and achieve maximum sensitivity. In other embodiments, the cutoff value can be determined, for example, as the mean plasma EBV DNA read percentage minus one standard deviation (SD), the mean minus two SDs, or the mean minus three SDs of NPC patients. In this example, the cutoff value was set to the mean plasma EBV DNA read percentage of all NPC patients minus three SDs. In yet other embodiments, the cutoff can be determined after logarithmic transformation of the percentage of plasma DNA fragments mapped to the EBV genome and then selected in a similar manner (e.g., using the mean, etc.). In still other embodiments, the cutoff can be determined using a receiver operating characteristic (ROC) curve by a nonparametric method, for example, including 100%, 95%, 90%, 85%, and 80% of the NPC patients under analysis. C. Combined analysis

[0162] These three techniques can be combined to improve accuracy. For example, each metric of each technique, which may include multiple metrics, e.g., multiple methylation levels, can be compared to a respective reference value to classify subjects (e.g., as can be done with a decision tree) or identify subjects corresponding to a particular section (quadrant) of a plot of training values. Clustering techniques, for example, as described above, can also be used.

[0163] We evaluated the combined value of plasma EBV DNA percentage (quantification) and size ratio analysis for NPC identification, as well as plasma EBV DNA percentage (quantification) analysis and methylation percentage for NPC identification, and then all three combined together.

[0164] Figure 25 is a plot of plasma EBV DNA read percentages and corresponding size ratio values ​​for NPC patients, transiently positive, and non-NPC subjects with persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. The same cutoff values ​​for EBV DNA size ratios and plasma EBV DNA read percentages specified in Figures 23 and 24 are shown as gray dotted lines. The oval highlights the quadrant 2510 that passed the combined analysis.

[0165] In this combined analysis, a plasma sample was considered positive if its sequence data simultaneously passed the cutoffs for both plasma EBV DNA percentage and size ratio analysis. Using the cutoffs defined above, the sensitivity, specificity, and positive predictive value for NPC detection were 90%, 88.5%, and 61.7%, respectively.

[0166] Figure 26 is a plot of plasma EBV DNA read percentages and corresponding methylation percentage values ​​for NPC patients, non-NPC subjects with transiently positive, and persistently positive plasma EBV DNA, according to an embodiment of the present disclosure. The same cutoff values ​​for EBV DNA methylation percentages and plasma EBV DNA read percentages within the 46 DMRs are shown with gray dotted lines, as specified in Figures 13 and 24, respectively. The oval highlights the quadrant 2610 that passed the combined analysis.

[0167] In this combined analysis, a plasma sample was considered positive if its sequencing data simultaneously passed the cutoffs for both plasma EBV DNA percentage and methylation percentage analysis (based on the DMRs specified in the figure). Using the cutoffs defined above, the sensitivity, specificity, and positive predictive value for NPC detection were 96.7%, 89.1%, and 64.6%, respectively.

[0168] Figures 27A and 27B show three-dimensional plots of plasma EBV DNA read percentages and corresponding size ratio and methylation percentage values ​​for NPC patients, non-NPC subjects with transiently positive, and persistently positive plasma EBV DNA, according to embodiments of the present disclosure. The combined value of all three parameters, plasma EBV DNA percentage (quantified), size ratio, and methylation percentage, for NPC identification was evaluated. Percentages were determined in the same manner as in Figure 24. Size ratios were determined in the same manner as in Figure 23. Methylation levels were then determined using the 46 DMRs used in Figure 13.

[0169] In Figure 27A, the purple surface 2710 shows a fitted 3D surface that distinguishes non-NPC subjects with transiently and persistently positive plasma EBV DNA from NPC patients in three-dimensional space. The use of a fitted surface demonstrates that the reference value for distinguishing subjects can be more complex than a constant value. For example, in Figure 26, the methylation percentage cutoff can vary depending on the proportion of sequence reads. Such flexibility in classification decisions can provide greater accuracy. Fitting can be selected to optimize various accuracy metrics, such as specificity, sensitivity, or an average of both. Such fitting can be performed using a support vector machine. Figure 27B shows the same data as Figure 27A, but without the surface 2710, and the axes are labeled in boxes around the data.

[0170] Figures 28A and 28B show receiver operating characteristic (ROC) curve analyses of various combinations of count-based, size-based, and methylation-based analyses according to embodiments of the present disclosure. Accuracy is for correctly classifying NPC and non-NPC subjects. Parameters (i.e., proportion, size ratio, and methylation level) were determined in the same manner as in Figures 27A and 27B. Varying cutoff values ​​resulted in changes in sensitivity and specificity. Figure 27A shows a comparison of the three techniques used individually. Area under the curve (AUC) values ​​are displayed. The AUC values ​​were 0.905, 0.942, and 0.979 for count-only, size-only, and methylation-only, respectively. Methylation yields the best results. Figure 28B shows a comparison of the three combined techniques. The AUC value for count and size is 0.97. The AUC value for count and methylation is 0.985. Using all three techniques yields the best accuracy at 0.989. VI. Other Virus Examples

[0171] Other viruses are also associated with cancer. For example, plasma human papillomavirus (HPV) is associated with head and neck squamous cell carcinoma (HNSCC). Hepatitis B virus (HBV) is also associated with hepatocellular carcinoma (HCC). The results below demonstrate that embodiments can use the methylation level(s) of other cell-free viral DNA to classify disease levels in a manner similar to that used for cell-free EBV DNA. A.HPV

[0172] Figure 29 shows the clinical staging of five cases of HPV-positive head and neck squamous cell carcinoma (HPV+ve HNSCC). The cases were staged according to the AJCC Cancer Staging Manual, 8th edition. The methylation profiles of plasma human papillomavirus (HPV) DNA reads from five patients with HPV-positive head and neck squamous cell carcinoma (HPV+ve HNSCC) were analyzed by targeted bisulfite sequencing of plasma DNA. All five patients had early-stage disease (stage I or II). All patients had detectable HPV DNA fragments in their plasma DNA samples.

[0173] For each clinical case, plasma DNA was extracted from 4 mL of plasma using a QIAamp DSP DNA Blood Mini Kit. In each case, all extracted DNA was used to prepare sequencing libraries using the TruSeq DNA PCR-Free Library Preparation Kit (Illumina). The adapter-ligated DNA products were subjected to two rounds of bisulfite treatment using the EpiTect Bisulfite Kit (Qiagen). 12–15 cycles of PCR amplification were performed on the bisulfite-converted samples using the KAPA HiFi HotStart Uracil+ReadyMix PCR Kit (Roche). The amplified products were then captured with the SeqCap-Epi system (Nimblegen) using custom-designed probes covering the viral and human genomic regions listed above (Figure 1).

[0174] After target capture, the captured products were enriched with 14 cycles of PCR to generate DNA libraries. The DNA libraries were sequenced on the NextSeq platform (Illumina). For each sequencing run, four to six samples with unique sample barcodes were sequenced using paired-end mode. For each DNA fragment, 75 nucleotides were sequenced from each of the two ends. After sequencing, the sequence reads were processed using Methyl-Pipe (Jiang et al. PLoS One 2014;9:e100360), a methylation data analysis pipeline, and mapped to artificially combined reference sequences including the entire human genome (hg19), EBV genome (AJ507799.2), HBV genome, and HPV genome. Sequenced reads that mapped to unique positions within the combined genome sequences (although mismatches may be tolerated in other embodiments) were used for downstream analysis.

[0175] Figure 30 shows plasma HPV DNA methylation profiles in individual patients with HPV-positive head and neck squamous cell carcinoma (HPV+ve HNSCC), according to an embodiment of the present disclosure. The HPV DNA methylation profiles were generated by targeted capture bisulfite sequencing of the plasma DNA of these patients. Figure 30 shows that HPV16 (HPV serotype 16) methylation can be detected in plasma HPV16 DNA molecules from HNSCC patients. The methylation density of all CpG sites throughout the HPV genome was determined.

[0176] The methylation density (MD) of unique loci in the entire viral genome in plasma can be calculated using the equation: MD = M / (M + U), where M is the number of methylated viral reads and U is the number of unmethylated viral reads at CpG sites within the locus in the entire viral genome. At the locus-specific level, these loci can be of any size and contain at least one CpG site (1 bp). If there are two or more CpG sites within a locus, M and U correspond to the total number of sites. These loci may or may not be associated with annotated viral genes.

[0177] Similar patterns of plasma HPV DNA methylation profiles were observed among different patients. Non-cancer subjects typically lack HPV. Two genomic regions, region 3001 and region 3002, were defined. Region-specific methylation densities in region 3001 were consistently higher than those in region 3002 in all five HPV+ve HNSCC cases. These similarities in patterns among subjects with the same disease state and differences in methylation profiles among subjects with different disease states can be analyzed at a global or locus-specific level. The methylation densities of such predefined regions can predict clinical manifestations, such as cancer stage, response to treatment, and risk of recurrence.

[0178] Figure 31 shows methylation levels of all CpG sites across the HPV genome in two patients with HPV+ve HNSCC, according to embodiments of the present disclosure. The first patient is shown in black and at the bottom, as shown in 3101. The second patient is gray and may appear larger, as shown in 3102. As shown, the two patients have apparent methylation levels at similar loci, with many loci having similar values ​​for methylation levels. By comparing methylation levels at global or locus-specific levels between cases, embodiments can identify subjects with HNSCC. B.HBV

[0179] Targeted bisulfite sequencing of plasma DNA was performed on nine patients with chronic hepatitis B virus infection and 10 patients with HCC. We also analyzed the methylation profiles of plasma HBV reads from these patients with chronic hepatitis B virus infection and HCC.

[0180] Figures 32A and 32B show the percentage of hepatitis B virus (HBV) DNA reads (plasma DNA reads mapped to the HBV genome) and the methylation percentage of all CpG sites in the entire HBV genome for nine patients with chronic hepatitis B virus infection (HBV) and ten patients with hepatocellular carcinoma (HCC), according to embodiments of the present disclosure.

[0181] In Figure 32A, the percentage of DNA fragments that aligned to the HBV genome relative to the human genome was determined. In this particular example, the number of cell-free DNA fragments that uniquely aligned to the HBV genome was divided by the number of cell-free DNA fragments that uniquely aligned to the combined reference genome, including human, HBV, HPV, and EBV genomes, listed in Figure 1. A higher mean percentage of HBV DNA reads was observed in HCC patients (mean = 0.006%) than in chronic HBV infection patients (mean = 0.03%), but statistical significance was not achieved (p = 0.07, Student's t-test). As can be seen visually, the means for HCC and HBV subjects are similar. Therefore, number-based techniques have relatively low predictive ability.

[0182] In Figure 32B, the HBV methylation percentage was determined as the global methylation level across all CpG sites in the HBV genome. It was observed that the methylation percentage of plasma HBV DNA in HCC patients (mean = 1.7%) was significantly higher than that in chronic hepatitis B virus infection patients (mean = 23%) (p = 0.03, Student's t-test). This separation indicates an improved ability to distinguish subjects with HBV but not HCC from those with HCC. Thus, embodiments can classify the level of HCC pathology (e.g., HCC or non-HCC). Therefore, embodiments can predict the risk of HCC based on the genome-wide methylation level of plasma HBV DNA in a sample.

[0183] Other embodiments may implement other types of methylation levels, e.g., as described herein, For example, classification may be based on methylation levels within differentially methylated regions with predefined criteria. VII. Methods using cell-free viral DNA methylation

[0184] As described above, embodiments can measure one or more methylation levels in samples of cell-free DNA, including cell-free viral DNA and cell-free genomic human DNA. Methylation levels can be used to classify the level of a pathology by analyzing the methylation level(s) of DNA from specific viruses associated with the pathology. Number-based and size-based techniques can also be used to complement methylation techniques.

[0185] As an example, the level of a condition can be whether the condition is present, the severity of the condition, the stage of the condition, the outlook for the condition, the response of the condition to treatment, or another measure of the severity or progression of the condition. As an example of cancer, the level of cancer can be whether the cancer is present, the stage of the cancer (e.g., early and late), the size of the tumor, the response of the cancer to treatment, or another measure of the severity or progression of the cancer.

[0186] For EBV, example conditions include infectious mononucleosis (IM), nasopharyngeal carcinoma (NPC), natural killer (NK)-T cell lymphoma, and subjects without these conditions who may exhibit a significant number of cell-free EBV DNA fragments in a sample. For HPV, example conditions include head and neck squamous cell carcinoma (HNSCC) and subjects with a significant amount of cell-free HPV DNA fragments but who do not have HNSCC. For HBV, example conditions include hepatocellular carcinoma (HCC) and subjects with a significant amount of cell-free HBV DNA fragments but who do not have HCC. A. Using Methylation Levels to Classify Disease States

[0187] 33 is a flow chart illustrating a method 3300 for analyzing a biological sample from an animal subject to determine a first condition classification, according to an embodiment of the present disclosure. The sample can include a mixture of subject DNA molecules and, in some cases, viral DNA molecules. Method 3300 can include clinical, laboratory, and in silico (computer) steps. Method 3300 can be performed as part of a subject screening, for example, to screen for cancer. Thus, the subject can be asymptomatic for the condition.

[0188] At block 3310, a biological sample is obtained from the subject. By way of example, the biological sample may be blood, plasma, serum, urine, saliva, sweat, tears, and sputum, as well as other examples provided herein. The biological sample may include a mixture of cell-free DNA molecules from the subject's genome and one or more other genomes. For example, the one or more other genomes may include a viral genome, such as an EBV, HPV, and / or HBV genome. In some embodiments (e.g., for blood), the biological sample may be purified for purposes of centrifugation of the blood to obtain a mixture of cell-free DNA molecules, e.g., plasma.

[0189] At block 3320, a plurality of cell-free DNA molecules are analyzed from the biological sample. Analysis of the cell-free DNA molecules may include identifying the location of the cell-free DNA molecules within a particular viral genome and determining whether the cell-free DNA molecules are methylated at one or more sites in the particular viral genome. Various numbers of cell-free DNA molecules (human and viral) can be analyzed, with various numbers (e.g., 10, 20, 30, 50, 100, 200, 500, or 1,000 or more) being identified as being from a particular viral genome, for example, at least 1,000.

[0190] The methylation status of a site in a sequence read can be obtained as described herein. For example, DNA molecules can be analyzed using sequence reads of the DNA molecules, in which case sequencing is methylation recognition. Other methylation recognition assays can also be used. Each sequence read can include the methylation status of a cell-free DNA molecule in a biological sample. The methylation status can include whether a particular cytosine residue is 5-methylcytosine or 5-hydroxymethylcytosine. Sequence reads can be obtained by various methods, including various sequencing technologies, PCR technologies (e.g., real-time or digital), arrays, and other suitable technologies for identifying the sequence of a fragment. Real-time PCR is an example of collectively analyzing a population of DNA, for example, as an intensity signal proportional to the number of methylated DNAs at a site. A sequence read can cover two or more sites, depending on the proximity of the two sites to each other and the length of the sequence read.

[0191] The analysis can be performed by receiving sequence reads from methylation-aware sequencing, and therefore, the analysis can be performed only on data previously obtained from DNA. In other embodiments, the analysis can include actual sequencing or other active steps that measure the properties of DNA molecules. Sequencing can be performed in various ways, for example, using massively parallel sequencing or next-generation sequencing, using single-molecule sequencing, and / or using double-stranded or single-stranded DNA sequencing library preparation protocols, as well as other techniques described herein. As part of the sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids.

[0192] The sequencing can be targeted sequencing, for example, as described herein. For example, a biological sample can be enriched for nucleic acid molecules from a virus. Enriching a biological sample for nucleic acid molecules derived from a virus can include using a capture probe that binds to a portion of the virus or the entire genome of the virus. Other embodiments can use primers specific to a particular locus of the virus. The biological sample can be enriched for nucleic acid molecules derived from a portion of the human genome, for example, an autosomal region. Figure 1 shows an example of such a capture probe. In other embodiments, the sequencing can include random sequencing.

[0193] After sequencing by the sequencing device, sequence reads can be received by a computer system that can be communicatively coupled to the sequencing device, for example, via wired or wireless communication or a removable storage device. In some embodiments, one or more sequence reads, including both ends of a nucleic acid fragment, can be received. The location of a DNA molecule can be determined by mapping (aligning) one or more sequence reads of the DNA molecule to a specific region, such as a respective portion of the human genome, e.g., a differentially methylated region (DMR). In one embodiment, if a read does not map to a region of interest, the read can be ignored. In other embodiments, a specific probe (e.g., after PCR or other amplification) can indicate location via a specific fluorescent color, etc. Identification can be such that the cell-free DNA molecule corresponds to one of a set of one or more sites; i.e., the specific site can be unknown, since all that is required is the amount of DNA methylated at one or more sites.

[0194] At block 3330, one or more mixture methylation levels are measured based on the amount of one or more of the plurality of cell-free DNA molecules methylated at a set of one or more sites of a particular viral genome. The mixture methylation level may be a methylation density or percentage (e.g., as described herein) of cell-free DNA molecules at a set of site(s) or a subset of site(s). For example, the methylation level may correspond to a methylation density determined based on the number of DNA molecules corresponding to the set of sites and the number methylated. The number may be determined based on alignment of the sequence reads to the viral genome in combination with the methylation state of a given sequence read at one or more sites.

[0195] The number of each DNA molecule that is methylated at a site can be determined for each set of sites. In one embodiment, the site is a CpG site, and can be only a specific CpG site selected using one or more criteria as described herein. The number of methylated DNA molecules is equivalent to determining the number of unmethylated DNA molecules when normalization is performed using the total number of DNA molecules analyzed at a specific site, for example, the total number of sequence reads. For example, an increase in the CpG methylation density of a certain region is equivalent to a decrease in the density of unmethylated CpG in the same region.

[0196] When one or more sets of sites include at least two sites, a single mixed methylation level can be determined for the at least two sites. For example, the methylation level can be calculated as the total methylation density of all cell-free DNA molecules in the first set. In another example, an individual methylation density can be calculated for each site or region(s) of one or more sites, thereby providing N (e.g., an integer greater than or equal to 2) mixed methylation levels as multidimensional points, e.g., as described in Sections IV.B-IV.E. The individual methylation densities can be combined to obtain a mixed methylation level, e.g., an average value of the individual methylation densities.

[0197] In other embodiments, distinct methylation levels can be retained for subsequent analysis, for example, using clustering and other techniques described herein. For example, multidimensional points (N mixed methylation levels) can be compared to N reference methylation levels to obtain N differences, which can be used to determine whether a subject belongs to one of at least two cohorts. Figures 18-20 show examples of such hierarchical clustering analysis. Regions can be predetermined, for example, as described above for predetermined regions having a size of 50 bases to 1,000 bases, with the size being the same or different between regions.

[0198] When two or more mixed methylation levels are determined, different levels may correspond to different subsets of the set of sites. For example, methylation levels can be determined for different regions, each of which may contain one or more sites. The regions can span the entire viral genome or correspond to only a portion, as can be done, for example, when specific regions are selected. Such regions can be selected to be differentially methylated according to one or more criteria, for example, as described herein. Such criteria can correspond to methylation levels in a cohort of subjects with the same condition within a specific range and potentially within a threshold difference from other cohorts of subjects. Thus, the criteria for a region or each site can include (1) differences in methylation levels between multiple subjects in the same cohort, and / or (2) differences in methylation levels between subjects in one cohort and subjects in another cohort.

[0199] At block 3340, the one or more mixture methylation levels are compared to one or more reference methylation levels determined from other subjects in at least two cohorts. The at least two cohorts can have different classifications associated with specific viral genomes, including a first disease state. Examples of such conditions, such as NPC, IM, lymphoma, and non-NPC conditions associated with transient or persistent positivity for the number of EBV DNA molecules, are shown above. Examples of cohorts and reference methylation levels are shown in Figures 8, 11, 13, 15, and 18-20.

[0200] The comparison can take various forms. For example, a separation value can be determined, such as a ratio or difference between the mixture methylation level and the reference methylation level. Various separation values ​​can be defined, including definitions that include ratios, differences, and functions of both. The comparison can further include comparing the separation value to a cutoff value to determine statistical significance. For example, the reference methylation level is the mean value of the cohort, and the difference between the mixture methylation level of the sample and the mean value of the cohort can be compared to a cutoff value, which can be determined based on the standard deviation of the methylation levels measured for the reference samples used in the cohort.

[0201] In some embodiments including multiple methylation levels and multiple reference levels, the multiple methylation levels correspond to multidimensional points of the sample (e.g., N levels forming a vector), while the multiple reference levels can correspond to an N-1 dimensional surface (e.g., a hyperplane), where the surface can be a closed surface, e.g., a sphere-like surface, where the data points correspond to the same disease state. As another example, comparison of the multiple methylation levels to the reference levels can be implemented by determining the distance from the multidimensional points of the sample to a representative (reference) multidimensional point of the reference subject. The reference multidimensional point can correspond to a single reference subject, e.g., patient AL038. As another example, the representative (reference) multidimensional point can be the centroid of a cluster of reference multidimensional points from subjects with the same disease state.

[0202] As another example, comparing the mixed methylation level(s) with the reference methylation level(s) can include inputting one or more mixed methylation levels into a machine learning model trained using one or more reference methylation levels determined from at least two cohorts of other subjects. For example, the reference levels can be methylation levels measured for other subjects, and the clustering model can be trained using the reference levels. For example, a centroid can be selected for a cluster of subjects corresponding to a particular cohort.

[0203] At block 3350, a first classification of whether the subject has the first condition is determined based on the comparison. The first classification can take various forms, such as a binary outcome or a probability value. In some embodiments, the first classification can provide a level of the first condition, such as tumor size, severity, or stage of cancer.

[0204] The different classifications of at least two cohorts can also include a second pathology, and by comparing with one or more reference levels, can determine whether the subject has a second pathology.For example, a single reference level can distinguish between IM and lymphoma, or between NPC and persistently positive subjects.

[0205] One or more mixed methylation levels can be compared to multiple reference methylation levels. As an example of using one methylation level but multiple reference methylation levels, different reference levels can distinguish between different disease states. For example, a first reference methylation level can determine a first classification of whether a subject has a first disease state (e.g., distinguishing between IM and no IM), and a second reference methylation level can determine a second classification of whether a subject has a second disease state (e.g., distinguishing between NPC and no NPC, as shown in Figures 8, 11, and 15). Thus, embodiments can use different reference levels to determine a second classification of whether a subject has a second disease state based on this comparison.

[0206] The method may further include treating the subject for the condition in response to classifying the subject as having the condition, thereby ameliorating the condition (e.g., eliminating or reducing the severity of the condition). If the condition is cancer, the treatment may include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, or precision medicine. Based on the determined level of the condition, a treatment plan may be developed to reduce the risk of harm to the subject. The method may further include treating the subject according to the treatment plan.

[0207] Biological samples can be obtained at various time points and analyzed independently or in conjunction with measurements and classifications at other time points. Examples of such time points include before and after cancer treatment (targeted therapy, immunotherapy, chemotherapy, surgery, etc.), different time points after cancer diagnosis, before and after cancer progression, before and after the occurrence of metastasis, before and after an increase in disease severity, or before and after the onset of complications. B. Using methylation levels in combination with size / number

[0208] As described in Section V, number-based and / or size-based techniques can be used in combination with methylation techniques. Such techniques can be implemented independently, e.g., each providing a separate classification. Each such independent classification may be required to provide the same result in order to provide a final classification of its results. In other embodiments, the reference values ​​for different techniques may depend on a metric from another technique. For example, a size reference value may depend on the methylation level measured for a given sample, as described above with respect to FIG. 27. In some embodiments, each metric (e.g., methylation level) may be a different element in a vector, thereby creating multidimensional data points from the metrics for a given sample. Each of the metrics may be determined from the same sample or separate samples, which may, for example, be obtained from the subject at approximately the same time.

[0209] In some embodiments, size-based techniques can be implemented as follows to analyze a subject's biological sample. The sample can be the same or a different sample as used for methylation analysis. The biological sample can include a mixture of cell-free DNA molecules from the subject's genome and one or more other genomes (e.g., viral genomes). The size and location of each of the multiple cell-free DNA molecules in the biological sample can be determined, e.g., as described herein. For example, both ends of the DNA molecule can be sequenced (e.g., to provide a single sequence read for the entire DNA molecule or a pair of sequence reads for both ends), and the sequence read(s) can be aligned to a reference genome to determine size. Thus, embodiments can measure the size of DNA molecules and identify the location of the DNA molecules within a specific viral genome. These cell-free DNA molecules can be the same as those used for methylation analysis, e.g., where methylation-aware sequencing is used. The sizes of the multiple DNA molecules can form a size distribution.

[0210] A size distribution statistic (such as a size ratio) can be determined. The statistic can be compared to reference size values ​​determined from other subjects in at least two cohorts, which can be the same two cohorts used for methylation analysis. The at least two cohorts can have different classifications (including a first pathology) associated with a particular viral genome. A size-based classification of whether the subject has the first pathology can be determined based on this comparison. The size-based classification and the methylation-based classification can be used together to provide a final classification. Examples of reference size values ​​are shown in Figures 22, 23, 25, and 27.

[0211] In some embodiments, the number-based technique can be implemented as follows to analyze a subject's biological sample. The sample can be the same or different from the sample used for methylation analysis. The biological sample can contain a mixture of cell-free DNA molecules from the subject's genome and one or more other genomes (e.g., viral genomes).

[0212] The amount of cell-free DNA molecules derived from a specific viral genome in a sample can be determined. In some embodiments, for each of a plurality of cell-free DNA molecules in a biological sample, whether the molecule is derived from a specific viral genome is determined, for example, by using sequencing or a probe, possibly with amplification such as PCR. For example, the location can be determined, for example, whether it is from the human genome or a specific viral genome. The location can be determined using a plurality of sequence reads obtained by sequencing a mixture of cell-free DNA. The amount of a plurality of sequence reads aligned to a specific viral genome can be determined. For example, the ratio of sequence reads aligned to the viral genome to the total number of sequence reads can be determined. The total number of sequence reads can be the sum of the sequence reads aligned to the reference genome corresponding to the virus and the sequence reads aligned to the human genome. Other ratios described herein can also be used, for example, the amount of reads from a specific viral genome divided by the amount of human reads.

[0213] The amount of sequence reads aligned to the reference genome can be compared to reference values ​​determined from other subjects in at least two cohorts, which can be the same two cohorts used for methylation and / or size analysis. The at least two cohorts can have different classifications (including a first pathology) associated with a particular viral genome. A count-based classification of whether the subject has the first pathology can be determined based on this comparison. The count-based classification and the methylation-based classification can be used together to provide a final classification. Examples of reference count values ​​are shown in Figures 24-27 and 32A. VIII. Exemplary Systems

[0214] FIG. 34 illustrates a system 3400 according to one embodiment of the present invention. The illustrated system includes a sample 3405, such as cell-free DNA molecules, in a sample holder 3410, which can contact an assay 3408 to provide a signal of a physical characteristic 3415. An example of a sample holder can be a flow cell, which includes assay probes and / or primers or a tube through which droplets (along with the droplets containing the assay) travel. The physical characteristic 3415, such as a fluorescence intensity value, from the sample is detected by a detector 3420. The detector 3420 can take measurements at intervals (e.g., periodic intervals) to obtain data points that constitute a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector to digital form multiple times. The sample holder 3410 and detector 3420 can form an assay device, such as a sequencing instrument, that performs sequencing according to embodiments described herein. The data signal 3425 is transmitted from the detector 3420 to a logic system 3430. The data signal 3425 may be stored in a local memory 3435 , an external memory 3440 , or a storage device 3445 .

[0215] Logic system 3430 may be or include a computer system, ASIC, microprocessor, etc. It may also include or be coupled to a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.). Logic system 3430 and other components may be part of a standalone or networked computer system, or may be directly attached to or incorporated into the thermal cycler device. Logic system 3430 may also include optimization software that executes in processor 3450. Logic system 3430 may include a computer-readable medium that stores instructions for controlling system 3400 to perform any of the methods described herein.

[0216] Any of the computer systems referred to herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG. 35 of computer system 10. In some embodiments, a computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, a computer system may include multiple computer devices, each of which is a subsystem and has internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0217] The subsystems shown in FIG. 35 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, storage device(s) 79, a monitor 76 coupled to a display adapter 82, and others. Peripherals and input / output (I / O) devices coupled to the I / O controller 71 may be connected to the computer system by any number of means known in the art, such as input / output (I / O) ports 77 (e.g., USB, FireWire®). For example, the I / O ports 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) may be used to connect the computer system 10 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of instructions from the system memory 72 or storage device(s) 79 (e.g., fixed disks such as hard drives or optical disks) and the exchange of information between the subsystems. The system memory 72 and / or storage device(s) 79 may embody a computer-readable medium. Another subsystem is a data collection device 85, such as a camera, microphone, and accelerometer. Any of the data mentioned herein may be output from one component to another or to a user.

[0218] A computer system may include multiple identical components or subsystems connected together, for example, by an external interface 81, by an internal interface, or through storage devices that may be connected or disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such an example, one computer may be considered a client and another computer a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0219] Aspects of the embodiments can be implemented using hardware circuitry in the form of control logic (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or computer software with general-purpose programmable processors in a modular or integrated manner. As used herein, a processor can include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked together, as well as dedicated hardware. Based on this disclosure and the teachings provided herein, those skilled in the art will recognize and understand other ways and / or methods for implementing embodiments of the present invention using hardware and combinations of hardware and software.

[0220] Any of the software components or functions described in this application may be implemented as software code executed by a processing device using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language, such as, for example, Perl or Python, using conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media (such as a hard drive or floppy disk), or optical media (such as a compact disc (CD) or DVD (digital versatile disc)), flash memory, and the like. The computer-readable medium may also be any combination of such storage or transmission devices.

[0221] Such programs may also be coded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks according to various protocols, including the Internet. Thus, computer-readable media may be created using data signals coded with such programs. Computer-readable media coded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, CD, or entire computer system), or may reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results described herein to a user.

[0222] Any of the methods described herein can be implemented in whole or in part using a computer system including one or more processors that can be configured to perform the steps. Accordingly, embodiments may be directed to a computer system configured to perform the steps of any of the methods described herein, potentially with different components performing each step or group of steps. While presented as numbered steps, steps of the methods herein can be performed simultaneously or at different times, or in different orders. In addition, portions of these steps can be combined with portions of other steps from other methods. Also, all or portions of a step may be optional. In addition, any of the steps of any of the methods can be performed using a module, unit, circuit, or other means of a system for performing these steps.

[0223] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of the invention, however, other embodiments of the invention may be directed to specific embodiments relating to each individual aspect or specific combinations of these individual aspects.

[0224] The above description of exemplary embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and many modifications and variations are possible in light of the above teachings.

[0225] All patents, patent applications, publications, and specifications mentioned herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art.

[0226] Appendix A [Table 1-1] Table 1-2 Table 1-3 Table 1-4 Table 1-5 Table 1-6 Table 1-7 Table 1-8 Table 1-9 Table 1-10 Table 1-11 Table 1-12 Table 1-13 Table 1-14 Table 1-15 Table 1-16 Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24

Claims

1. 1. A method for analyzing a biological sample from an animal subject, wherein the biological sample comprises a mixture of cell-free DNA molecules from the subject's genome and one or more other genomes, the method comprising: analyzing a plurality of cell-free DNA molecules from the biological sample, wherein analyzing one of the plurality of cell-free DNA molecules comprises: identifying the location of the cell-free DNA molecule in a particular viral genome; determining whether the cell-free DNA molecules are methylated at one or more sites of the specific viral genome; measuring one or more mixture methylation levels based on the amount of one or more of the plurality of cell-free DNA molecules methylated at the set of one or more sites of the particular viral genome; comparing the one or more mixed methylation levels to one or more reference methylation levels determined from at least two cohorts of other subjects, the at least two cohorts having different classifications associated with the particular viral genome, the different classifications including a first disease state; and determining a first classification of whether the subject has the first condition based on the comparison.

2. wherein the different classifications of the at least two cohorts further comprise a second disease state, and the method further comprises:

10. The method of claim 1, further comprising determining a second classification of whether the subject has the second condition based on the comparison.

3. 3. The method of claim 2, wherein the one or more mixed methylation levels are compared to a plurality of reference methylation levels, including a first reference methylation level and a second reference methylation level, wherein the first reference methylation level is used to determine the first classification of whether the subject has the first pathological condition, and the second reference methylation level is used to determine the second classification of whether the subject has the second pathological condition.

4. 4. The method of claim 3, wherein the specific viral genome is that of Epstein-Barr virus, the subject is a human, the first pathology is nasopharyngeal carcinoma, and the second pathology is infectious mononucleosis.

5. 10. The method of claim 1, wherein the first classification is that the subject does not have the first condition.

6. The method of claim 1 , wherein determining the first classification comprises determining a level of the first condition.

7. 2. The method of claim 1, wherein the specific viral genome is that of the Epstein-Barr virus, the subject is a human, and the first condition is nasopharyngeal carcinoma.

8. 2. The method of claim 1, wherein the set of one or more sites comprises at least two sites, and the one or more mixed methylation levels is a mixed methylation level determined across the at least two sites.

9. the one or more mixed methylation levels comprise N mixed methylation levels, where N is an integer greater than 1, the one or more sets of sites comprise at least two sites, and the comparison is measuring the difference between the N mixture methylation levels and the N reference methylation levels; and using the difference to determine whether the subject belongs to one of the at least two cohorts.

10. 10. The method of claim 9, wherein using the difference to determine whether the subject belongs to one of the at least two cohorts comprises performing a hierarchical clustering analysis.

11. 10. The method of claim 9, wherein each of the N mixed methylation levels is measured for one of a plurality of predetermined regions.

12. 12. The method of claim 11, wherein the plurality of predetermined regions are of the same size and span the specific viral genome, and the same size is between 50 bases and 1,000 bases.

13. 12. The method of claim 11 , wherein each of the plurality of predetermined regions satisfies one or more criteria including: (1) a difference in methylation levels between multiple subjects of the same cohort; and / or (2) a difference in methylation levels between subjects of one cohort and subjects of another cohort.

14. 2. The method of claim 1, wherein the set of one or more sites is present in multiple regions that each meet one or more criteria including: (1) a difference in methylation levels between multiple subjects of the same cohort; and / or (2) a difference in methylation levels between subjects of one cohort and subjects of another cohort.

15. 2. The method of claim 1, wherein the set of one or more sites meets one or more criteria including: (1) a difference in methylation levels between multiple subjects of the same cohort; and / or (2) a difference in methylation levels between subjects of one cohort and subjects of another cohort.

16. comparing the one or more mixed methylation levels to the one or more reference methylation levels determined from at least two cohorts of other subjects; 2. The method of claim 1, comprising inputting the one or more mixture methylation levels into a machine learning model trained using the one or more reference methylation levels determined from the at least two cohorts of other subjects.

17. For each site of the set of one or more sites:

10. The method of claim 1, further comprising determining the number of respective DNA molecules that are methylated at said sites, thereby determining the amount of said one or more of said plurality of cell-free DNA molecules methylated at said set of one or more sites of said particular viral genome.

18. performing methylation-aware sequencing of the plurality of cell-free DNA molecules to obtain sequence reads; 18. The method of Claim 17, further comprising aligning the sequence reads to the particular viral genome and determining the number of the respective DNA molecules that are methylated at each site in the set of one or more sites.

19. 10. The method of claim 1, further comprising performing a methylation recognition assay of the plurality of cell-free DNA molecules as part of determining the locations of the plurality of cell-free DNA molecules and whether the plurality of cell-free DNA molecules are methylated at the set of one or more sites.

20. 2. The method of claim 1, wherein identifying the location of the cell-free DNA molecule comprises determining that the location corresponds to one of the set of one or more sites.

21. 10. The method of claim 1, wherein the group of cell-free DNA molecules is collectively analyzed to determine the amount of one or more of the cell-free DNA molecules that are methylated at the set of one or more sites in the specific viral genome.

22. 2. The method of claim 1, wherein the plurality of cell-free DNA molecules comprises at least 10 cell-free DNA molecules located in the specific viral genome.

23. 2. The method of claim 1, wherein the specific viral genome corresponds to Epstein-Barr virus, human papillomavirus, or hepatitis B virus.

24. For each of a set of cell-free DNA molecules in the sample, measuring the size of the cell-free DNA molecules; identifying the location of the cell-free DNA molecules in the particular viral genome, wherein the sizes of the set of cell-free DNA molecules form a size distribution, and the sample is the biological sample or a different sample containing a mixture of cell-free DNA molecules from the subject's genome and the one or more other genomes; determining statistics of said size distribution; comparing said statistical value to reference size values ​​determined from said at least two cohorts of other subjects; determining a second classification of whether the subject has the first condition based on the comparison of the statistical value to the reference size value; and The method of claim 1 , further comprising: determining a final classification using the first classification and the second classification.

25. determining the amount of cell-free DNA molecules derived from the particular viral genome in a sample, wherein the sample is the biological sample or a different sample comprising a mixture of cell-free DNA molecules from the genome of the subject and the one or more other genomes; comparing said amount to reference values ​​determined from said at least two cohorts of other subjects; determining a second classification of whether the subject has the first condition based on the comparison of the amount to the reference value; The method of claim 1 , further comprising: determining a final classification using the first classification and the second classification.

26. 10. The method of claim 1, further comprising, in response to the first classification of the subject as having the first condition, providing the subject with a treatment that improves the first condition.

27. A computer product comprising a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any of the methods described above.

28. 1. A system comprising:

28. A computer product according to claim 27; one or more processors for executing instructions stored on the computer-readable medium.

29. A system comprising means for performing any of the above-mentioned methods.

30. A system configured to perform any of the above-mentioned methods.

31. A system comprising modules for performing each of the steps of any of the above-mentioned methods.