Stratification of risk of virus associated cancers

By analyzing cell-free nucleic acid molecules from pathogens in biological samples, the method addresses the limitations of current diagnostic techniques for pathogen-related disorders, offering a non-invasive risk assessment that informs appropriate screening frequencies and may lead to early detection.

JP2025084804APending Publication Date: 2025-06-03GRAIL INC

Patent Information

Application Number
JP2025024062
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-15
Filing Date
2025-02-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Current diagnostic methods for pathogen-related disorders, such as nasopharyngeal carcinoma associated with Epstein-Barr virus, are invasive or have high false positive rates, limiting their effectiveness in early detection and risk stratification.

Method used

A method involving the analysis of cell-free nucleic acid molecules from pathogens in biological samples, which includes determining characteristics like genetic abundance, methylation status, and mutation patterns to assess the risk of developing a pathogen-associated disorder. This method involves performing assays at specific time points, with the interval between assays inversely correlated with the risk level.

Benefits of technology

The method provides a non-invasive approach for assessing the risk of pathogen-related disorders, enabling appropriate screening frequencies and potentially leading to early detection and improved patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084804000001_ABST
    Figure 2025084804000001_ABST
Patent Text Reader

Abstract

To provide a method of screening a pathogen-associated disorder in a subject, a method of prognosticating a pathogen-associated disorder in a subject, and a system therefor.SOLUTION: Provided herein are methods and systems for stratifying risk for a subject to develop a pathogen-associated disorder based on analysis of cell-free nucleic acid molecules from a biological sample of the subject. In various examples, screening frequency is determined based on the risk analysis. Also provided herein are methods and systems for analyzing variant patterns of a pathogen genome in cell-free nucleic acid molecules.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefit of U.S. Provisional Application No. 62 / 961,517, filed Jan. 15, 2020, and U.S. Provisional Application No. 62 / 828,224, filed Apr. 2, 2019, each of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Many diseases and conditions may be associated with infections by pathogens such as viruses. Nasopharyngeal carcinoma (NPC) is one of the most prevalent cancers in southern China and Southeast Asia, and the etiology of NPC is likely closely associated with Epstein-Barr virus (EBV) infection. In regions with a high incidence of NPC, EBV genomes will likely be latent in almost all NPC tumors. Based on the close relationship between EBV and NPC, plasma EBV DNA has been developed as a biomarker for NPC. Using real-time polymerase chain reaction (PCR) analysis, the detection of plasma EBV DNA has been shown to have 95% sensitivity and 93% specificity for the detection of NPC (Lo et al. . Cancer Res. 1999; 59:1188-91). Based on the analysis of cell-free nucleic acid molecules from pathogens in biological samples, there is great clinical benefit in developing non-invasive or minimally invasive diagnostic assays to stratify the risk of these pathogen-related disorders.

Summary of the Invention

[0003] In one aspect, provided herein is a pathogen in a subject​​​​​ a method of screening for a related disorder, the method comprising: determining a characteristic of a cell-free nucleic acid molecule from the pathogen; receiving data from a first assay performed at said point, These cell-free nucleic acid molecules are characterized by their abundance, methylation status, variant patterns, fragment size, and or relative to the cell-free nucleic acid molecules from the subject in said biological sample. and the characteristic comprises a genetic abundance that is indicative of a risk that the subject will develop the pathogen-associated disorder. receiving a signal indicating the pathogen-associated disorder in the subject based on the characteristic; determining a second time point at which a second assay is performed to screen for toxicity; wherein the interval between the first time point and the second time point is inversely correlated with the risk. and determining whether to

[0004] In one aspect, provided herein is a method for treating a pathogen-associated disorder in a subject. 2. A method of prognostic diagnosis, the method comprising: determining a characteristic of acellular nucleic acid molecules from the pathogen in the biological sample of the subject. receiving data from a first assay, the data including: The properties of the acid molecule may be determined by the amount, methylation state, mutation pattern, fragment size, or the like. The relative abundance of the nucleic acid molecules in the received sample compared to cell-free nucleic acid molecules in a biological sample from the subject. determining a characteristic of the pathogen-derived cell-free nucleic acid molecule, as well as the age, pre-existing condition, and other characteristics of the subject; the subject's smoking habit, the subject's family history of a pathogen-associated disorder, the subject's genotypic factors, the subject based on one or more factors of the subject's ethnicity, or the subject's dietary history. including the step of a tester creating a report indicating the risk of developing a pathogen-related disorder .

[0005] In certain cases, the results of the first assay do not result in a result such as medical treatment of a subject with a pathogen-related disorder. In certain cases, medical treatment includes treatment with a therapeutic agent, radiation therapy or surgical treatment. In certain cases, the subject is diagnosed as not having a pathogen-related disorder prior to a determination at a second time point by a clinical diagnostic test with a false positive rate of less than 1%. In certain cases the clinical diagnostic test includes a physical examination, an invasive biopsy, an endoscopy, magnetic resonance imaging, positron emission tomography computed tomography, or X-ray imaging. In certain cases, the clinical diagnostic test includes an invasive biopsy including histological analysis, cytological analysis, or cellular nucleic acid analysis . In certain cases, the interval is at least about 2 months, 4 months, 6 months, 8 months, 10 months , or 12 months. In certain cases, the interval is at least about 12 months . In certain cases, the method further includes performing the first assay. In certain cases performing the first assay includes: (i) obtaining a first biological sample from the subject; and (ii) measuring a first amount of cell-free nucleic acid molecules from the pathogen in the first biological sample

[0006] . In certain cases, the measuring the first amount includes measuring the copy number of the cell-free nucleic acid molecules from the pathogen in the first biological sample . In certain cases, the measuring includes polymerase chain reaction (PCR) . In certain cases, the measuring includes quantitative PCR (qPCR). In certain cases . ​​​The first amount includes measuring the percentage of the first of the cell-free nucleic acid molecules from the pathogen in the first biological sample. In certain cases, the first assay is:( (iii) If the first amount exceeds a threshold, obtaining a second biological sample from the subject and further measuring a second amount of cell-free nucleic acid molecules from the pathogen in the second biological sample. In certain cases, the second biological sample is obtained approximately 4 weeks after the first biological sample. In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status iii) When the first amount exceeds a threshold, obtaining a second biological sample from the subject and measuring a second amount of cell-free nucleic acid molecules from the pathogen in the second biological sample. In certain cases, the second biological sample is obtained approximately 4 weeks after the first biological sample. In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status obtaining a second biological sample from the subject and measuring a second amount of cell-free nucleic acid molecules from the pathogen in the second biological sample. In certain cases, the second biological sample is obtained approximately 4 weeks after the first biological sample. In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the second biological sample is obtained approximately 4 weeks after the first biological sample. In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the second biological sample is obtained approximately 4 weeks after the first biological sample. In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is shorter when both the first amount and the second copy number exceed the threshold compared to the interval when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is longer when the first amount is below the threshold compared to the interval when the first amount exceeds the threshold. In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is about 1 year when both the first amount and the second amount exceed the threshold. In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is about 2 years when the second amount is below the threshold. In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the interval between the first time point and the second time point is about 4 years when the first amount is below the threshold. In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the first assay includes determining the methylation status of the cell-free nucleic acid molecules from the pathogen in the biological sample. In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status In certain cases, the determination of the methylation status includes treating the cell-free nucleic acid molecules in the biological sample with a methylation-sensitive restriction enzyme or bisulfite. In certain cases, the determination of the methylation status The determination includes performing methylation-aware sequencing of cell-free nucleic acids in the biological sample of the subject. In one case, the methylation-aware sequencing includes bisulfite conversion of unmethylated cytosine to uracil. In one case, the methylation-aware sequencing includes treatment with methylation-sensitive restriction enzymes. In one case, the first assay includes determining the fragment size distribution of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the fragment size distribution includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the fragment size of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the first assay includes determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the mutation pattern includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the mutation pattern of the cell-free nucleic acid molecules from the pathogen includes single nucleotide variations. In one case, the identifying of the mutation pattern includes comparing sequence reads mapped to the reference genome of the pathogen with a disorder-related reference genome of the pathogen. In one case, the first assay includes determining the fragment size distribution of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the fragment size distribution includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the fragment size of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the methylation-aware sequencing includes treatment with methylation-sensitive restriction enzymes. In one case, the first assay includes determining the fragment size distribution of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the fragment size distribution includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the fragment size of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the first assay includes determining the fragment size distribution of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the fragment size distribution includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the fragment size of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the first assay includes determining the fragment size distribution of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the fragment size distribution includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the fragment size of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen.

[0007] In one case, the first assay includes determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the determination of the mutation pattern includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the determination of the mutation pattern includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the determination of the mutation pattern includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the determination of the mutation pattern includes performing sequencing of the cell-free nucleic acid molecules in the biological sample and determining the mutation pattern of the cell-free nucleic acid molecules from the pathogen in the biological sample based on sequence reads mapped to the reference genome of the pathogen. In one case, the mutation pattern of the cell-free nucleic acid molecules from the pathogen includes single nucleotide variations. In one case, the identifying of the mutation pattern includes comparing sequence reads mapped to the reference genome of the pathogen with a disorder-related reference genome of the pathogen. In one case, the identifying of the mutation pattern includes comparing sequence reads mapped to the reference genome of the pathogen with a disorder-related reference genome of the pathogen. includes determining a similarity level therebetween. In one case, the disorder-related reference genome of the pathogen comprises the genome of the pathogen identified in the diseased tissue. In one case, the determination of the similarity level comprises: separating the reference genome of the pathogen into a plurality of bins; and determining a similarity index for each of the plurality of bins with respect to the disorder-related reference genome of the pathogen, wherein the similarity index correlates with the proportion of mutant sites in each of the bins in which at least one of the sequence reads mapped to the reference genome of the pathogen has the same nucleotide variant as the disorder-related reference genome of the pathogen, and determining. In one case, the disorder-related reference genome of the pathogen comprises a plurality of disorder-related reference genomes of the pathogen, and the determination of the similarity level comprises: determining a similarity index for each of the plurality of bins with respect to each of the plurality of disorder-related reference genomes of the pathogen; determining a bin score for each of the plurality of bins based on the proportion of the plurality of disorder-related reference genomes for which the respective similarity index within each of the respective bins exceeds a cut-off value. In one case, the lengths of the plurality of bins are each about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bp. In one case, the first assay comprises determining the methylation state, the fragment size distribution, or the mutation pattern of cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the method further comprises cell-free nuclei from the pathogen in the biological sample from the pathogen in the biological sample respectively, about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bp. In one case, the first assay comprises determining the methylation state, the fragment size distribution, or the mutation pattern of cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the first assay comprises determining the methylation state, the fragment size distribution, or the mutation pattern of cell-free nucleic acid molecules from the pathogen in the biological sample. In one case, the first assay comprises determining the methylation state, the fragment size distribution, or the mutation pattern of cell-free nucleic acid molecules from the pathogen in the biological sample.

[0008] In one case, the method further comprises cell-free nuclei from the pathogen in the biological sample Using a classifier applied to data input, including characteristics of acid molecules, to calculate a risk score for the subject to develop the pathogen-related disorder, wherein the classifier is configured to apply a function to the data input including characteristics of cell-free nucleic acid molecules from a pathogen in the biological sample, and to generate an output including the risk score for evaluating the risk of the subject developing the disorder. In one case, the classifier is trained with a labeled dataset. In one case, the method further includes performing the second assay at the second time point. In one case, the second assay is the same as the first assay. In one case, the second assay includes an assay of cell-free nucleic acid molecules from the subject, an invasive biopsy of the subject, an endoscopy of the subject, or a magnetic resonance imaging examination of the subject. In one aspect, provided herein is a method for analyzing nucleic acid molecules from a biological sample of a subject, the method comprising: obtaining, in a computer system, sequence reads of cell-free nucleic acid molecules from the biological sample of the subject, wherein the biological sample includes cell-free nucleic acid molecules from the subject and potentially from a pathogen; aligning, in the computer system, the sequence reads of the cell-free nucleic acid molecules to a reference genome of the pathogen; and identifying, in the computer system, a mutation pattern of cell-free nucleic acid molecules from the pathogen, wherein the mutation pattern is a plurality of mutation sites on the reference genome of the pathogen.

[0009]

[0010] ​​​​​​​​​​​​​​​​ For each of the to, the nuc of the sequence read mapped to the reference genome of the pathogen Characterize the leotide variants, and the plurality of mutation sites span the reference genome of the pathogen Including at least 30 sites, and the mutation pattern indicates the state of the pathogen-related disorder or its risk in the subject Identifying steps, including.

[0011] In some cases, the plurality of mutation sites span the reference genome of the pathogen, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 40 0, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 Sites are included. In some cases, the plurality of mutation sites span the reference genome of the pathogen Including the plurality of mutation sites including at least 600 sites. In some cases The plurality of mutation sites include the plurality of mutation sites including about 660 sites spanning the reference genome of the pathogen Including the plurality of mutation sites including at least 1000 sites spanning the reference genome of the pathogen. In some cases Including the plurality of mutation sites including at least 1000 sites spanning the reference genome of the pathogen In some cases, the plurality of mutation sites span the reference genome of the pathogen, about Including 1100 sites. In some cases, the plurality of mutation sites are the references of the pathogen The sequence reads mapped to the genome consist of all sites having nuc leotide mutations different from the reference genome of the pathogen In some cases, the sequence reads The sequence reads are aligned to a reference genome of the pathogen and There are 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs between the pathogen and the reference genome. In some cases, the sequence linker is configured to tolerate a maximum mismatch of 100 bp. and aligning the sequence reads to a reference genome of the pathogen. and configured to allow a maximum mismatch of 2 bases between the pathogen and a reference genome. In some cases, the method further comprises: mapping to a reference genome of the pathogen. a pathogen-associated disorder in the subject based on a mutation pattern of the sequence reads. In some cases, this includes diagnosing, prognosing, or monitoring the In some cases, the mutation pattern of the cell-free nucleic acid molecule from the pathogen comprises a single base mutation. The identification of the mutation pattern is: determining a level of similarity between the detection read and a disorder-associated reference genome of the pathogen; In some cases, the disease-associated reference genome of the pathogen includes the disease-associated reference genome of the pathogen identified in the diseased tissue. In some cases, the determination of the level of similarity is based on the genome of the pathogen. Separating the genome into a plurality of bins; and detecting the pathogen's pathogenicity against a disease-associated reference genome. determining a similarity index for each of a plurality of bins, said similarity index being At least one of the sequence reads mapped to a reference genome of the disease entity is A mutation sample within each of the bins that has the same nucleotide variation as the disease-associated reference genome of the original genome. In some cases, determining a proportion of the pathogen that is impaired by the pathogen. The related reference genome comprises a plurality of disorder-associated reference genomes of the pathogen, and the similarity level is Determining the bins: for each of the plurality of disorder-related reference genomes of the pathogen, determining a respective similarity index for each of the plurality of bins and; based on the ratio of the plurality of disorder-related reference genomes for which the respective similarity index within each of the respective bins exceeds a cut-off value, determining a bin score for each of the plurality of bins, including. In one case, the cut-off value is about 0.9. In one case, the lengths of the plurality of bins are each about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1 000 bp. In one case, the method further comprises: calculating a risk score for the subject to develop a disorder related to the pathogen using a classifier applied to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, wherein the classifier is configured to apply a function to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, and generating an output including the risk score for evaluating the risk of the subject developing a disorder including calculating. In one case, the classifier is trained on a labeled data set. In one case, the classifier is a Naive Bayes model, logistic regression, random forest, decision tree, gradient boosting tree, neural network, deep learning, linear / kernel support vector machine (SVM), linear / non-linear regression, or other machine learning algorithms. In one case, the method further comprises: determining a bin score for each of the plurality of bins based on the ratio of the plurality of disorder-related reference genomes for which the respective similarity index within each of the respective bins exceeds a cut-off value, including. In one case, the cut-off value is about 0.9. In one case, the lengths of the plurality of bins are each about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1 000 bp. In one case, the method further comprises: calculating a risk score for the subject to develop a disorder related to the pathogen using a classifier applied to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, wherein the classifier is configured to apply a function to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, and generating an output including the risk score for evaluating the risk of the subject developing a disorder including calculating. In one case, the classifier is trained on a labeled data set. In one case, the classifier is a Naive Bayes model, logistic regression, random forest, decision tree, gradient boosting tree, neural network, deep learning, linear / kernel support vector machine (SVM), linear / non-linear regression, or other machine learning algorithms. In one case, the method further comprises: determining a bin score for each of the plurality of bins based on the ratio of the plurality of disorder-related reference genomes for which the respective similarity index within each of the respective bins exceeds a cut-off value, including. In one case, the cut-off value is about 0.9. In one case, the lengths of the plurality of bins are each about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1 000 bp. In one case, the method further comprises: calculating a risk score for the subject to develop a disorder related to the pathogen using a classifier applied to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, wherein the classifier is configured to apply a function to data input including a mutation pattern of cell-free nucleic acid molecules from the pathogen, and generating an output including the risk score for evaluating the risk of the subject developing a disorder including calculating. In one case, the classifier is trained on a labeled data set. In one case, the classifier is a Naive Bayes model, logistic regression, random forest, decision tree, gradient boosting tree, neural network, deep learning, linear / kernel support vector machine (SVM), linear / non-linear regression, or other machine learning algorithms. In one case, the method further comprises: determining a bin score for each of the plurality of bins based on the ratio of the plurality of disorder-related reference genomes for which the respective similarity index within each of the respective bins exceeds a cut-off value, including. In one case, the cut-off value is about 0.9. In one case, the lengths of the plurality of bins are each about or a mathematical model using linear discriminative analysis.

[0012] In certain cases, the pathogen is a virus. In certain cases, the virus is Epstein-Barr virus (EBV). In certain cases, the pathogen-related disorder includes nasopharyngeal carcinoma, NK cell lymphoma, Burkitt's lymphoma, post-transplant lymphoproliferative disorders, or Hodgkin's lymphoma. In certain cases, the mutation pattern of the cell-free nucleic acid molecule from the pathogen maps to the reference genome of the pathogen at each of at least 30, 40, 50, 100, 1 50, 200, 250, 300, 350, 400, 450, 500, 550, or 60 0 sites selected from the genomic sites described in Table 6 related to the EBV reference genome (AJ507799.2), characterizing the nucleotide variant of the sequence code. In certain cases, the plurality of mutation sites described above include the genomic sites described in Table 6 related to the EBV reference genome (AJ507799.2). In certain cases, the mutation pattern of the cell-free nucleic acid molecule from the pathogen is randomly selected from the genomic sites described in Table 6 related to the EBV reference genome (AJ507799.2), characterizing the nucleotide variant of the sequence At least 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or 600 sites randomly selected from the genomic sites described in Table 6 related to the reference genome (AJ507799.2), characterize the nucleotide variants of the sequence reads mapped to the reference genome of the pathogen at each of the plurality of mutant sites. In certain cases, the virus is a human papilloma virus (HPV). In certain cases, the pathogen-related disorder includes cervical cancer, oropharyngeal cancer, or head and neck cancer. In certain cases, the virus is a hepatitis B virus (HBV). In certain cases, the pathogen-related disorder includes cirrhosis or hepatocellular carcinoma (HCC). In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof.

[0013] In certain cases, the virus is a human papilloma virus (HPV). In certain cases, the pathogen-related disorder includes cervical cancer, oropharyngeal cancer, or head and neck cancer. In certain cases, the virus is a hepatitis B virus (HBV). In certain cases, the pathogen-related disorder includes cirrhosis or hepatocellular carcinoma (HCC). In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the virus is a human papilloma virus (HPV). In certain cases, the pathogen-related disorder includes cervical cancer, oropharyngeal cancer, or head and neck cancer. In certain cases, the virus is a hepatitis B virus (HBV). In certain cases, the pathogen-related disorder includes cirrhosis or hepatocellular carcinoma (HCC). In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the virus is a human papilloma virus (HPV). In certain cases, the pathogen-related disorder includes cervical cancer, oropharyngeal cancer, or head and neck cancer. In certain cases, the virus is a hepatitis B virus (HBV). In certain cases, the pathogen-related disorder includes cirrhosis or hepatocellular carcinoma (HCC). In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the virus is a human papilloma virus (HPV). In certain cases, the pathogen-related disorder includes cervical cancer, oropharyngeal cancer, or head and neck cancer. In certain cases, the virus is a hepatitis B virus (HBV). In certain cases, the pathogen-related disorder includes cirrhosis or hepatocellular carcinoma (HCC). In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the mutation pattern indicates the state of the pathogen-related disorder in the subject, and the state of the pathogenic-related disorder includes the presence of the pathogenic-related disorder in the subject, the amount of tumor tissue in the subject, the size of the tumor tissue in the subject, the stage of the tumor in the subject, the tumor burden in the subject, or the presence of tumor metastasis in the subject. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. In certain cases, the biological sample is selected from the group consisting of whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, vaginal flushing fluid, saliva, oral rinse fluid, nasal flushing fluid, nasal brush sample, and combinations thereof. ​

[0014] In one aspect, provided herein is machine-executable code for implementing any of the above methods by execution by one or more computer processors and a non-transitory computer-readable medium containing the same.

[0015] In one aspect, provided herein is a computer product including a non-transitory computer-readable medium storing a plurality of instructions for controlling a computer system to perform any of the above operations.

[0016] In one aspect, provided herein is a system including: the computer product described herein; and one or more processors for executing instructions stored on the computer-readable medium.

[0017] In one aspect, provided herein is a system including means for performing any of the above methods.

[0018] In one aspect, provided herein is a system configured to perform any of the above methods.

[0019] In one aspect, provided herein is a system including modules for performing each of the steps of any of the above methods.

[0020] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent, or In the same degree as the case where it is specifically and individually shown that the patent application is incorporated by reference, the reference is incorporated into this specification. is incorporated into this specification by reference.

[0021] The novel features described in this specification are specifically described in the appended claims. A better understanding of the features and advantages described in this specification can be obtained by referring to the following detailed description of exemplary embodiments in which the principles described in this specification are utilized, and the accompanying drawings thereof. is incorporated into this specification by reference. is obtained by referring to the following detailed description of exemplary embodiments that illustrate the principles described in this specification, and the accompanying drawings thereof.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

[0023] Summary In an aspect, provided herein are methods and systems for screening for pathogen-related disorders in a subject. for learning. The methods and systems are based on the characteristics of cell-free nucleic acid molecules from pathogens in a biological sample from the subject, the subject from the biological sample from the subject, based on the characteristics of cell-free nucleic acid molecules from pathogens in the biological sample from the subject, the subject can provide an assessment of the risk that the subject develops the pathogen-related disorder. Among other things, risk prediction enables the determination of an appropriate screening frequency. Appropriate and timely follow-up screening not only saves the cost of the subject, but also enables early detection of the disorder. For example, a shift in the stage distribution to the early stage of EBV-NPC may result in a significant improvement in the overall survival period without progression of NPC patients.

[0024] The risk that the subject develops the pathogen-related disorder can refer to the possibility that the subject has a tendency to develop the pathogen-related disorder. In one case, the risk described in this specification can refer to the possibility that the pathogen-related disorder develops into a state clinically detectable in the subject at a future time point (a "clinically detectable disorder"). In one case, the subject is screened at an initial time point by a screening assay that tests for cell-free nucleic acid molecules from the pathogen in a biological sample from the subject, and the subject is diagnosed as not having a clinically detectable pathogen-related disorder at the initial time point, but the characteristics of the cell-free nucleic acid molecules from the pathogen in the biological sample from the subject can indicate the risk that the subject has a clinically detectable disorder at a future

[0025] A clinically detectable disorder can refer to a disorder that defines pathological symptoms that can be detected through one or more well-established clinical diagnostic tests. In one case, the well-established clinical diagnostic tests include medical tests / assays with a low false positive detection rate for the pathogen-related disorder, and the ratio is, for example, 30%, 20%, 10%, 8%, 7%, 6 %, 5%, 4%, 3%, 2.5%, 2%, 1%, 0.8%, 0.5%, 0.25%, 0. 15%, less than 0.1%, 0.08%, 0.05%, 0.02%, 0.01%, 0.005 %, 0.002%, 0.001% or less. The well - established clinical diagnosis test includes a medical test / assay that can detect the pathogen - related disorder with high sensitivity, and the proportion is, for example, at least 30%, 40%, 50%, 60%, 70%, 80% including, for example, at least 30%, 40%, 50%, 60%, 70%, 80% , 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%, or 100%. In some cases, the pathogen - related disorder is a pathogen - related proliferative disorder such as cancer, and the cancer is diagnosed by one or more invasive biopsies followed by histological or other examinations (such as cell examinations such as tissue analysis, cell DNA or protein analysis) of the biopsy tissue, for example, imaging examinations such as X - ray, magnetic resonance imaging (MRI), positron emission tomography (PET) of the biopsy tissue, for example, histological or other examinations (such as tissue analysis, cell DNA or protein analysis) of the biopsy tissue, for example, imaging examinations such as X - ray, magnetic resonance imaging (MRI), positron emission tomography (PET) histological or other examinations (such as tissue analysis, cell DNA or protein analysis) of the biopsy tissue, for example, imaging examinations such as X - ray, magnetic resonance imaging (MRI), positron emission tomography (PET) T), for example, imaging examinations such as X - ray, magnetic resonance imaging (MRI), positron emission tomography (PET) T), or computed tomography (CT), or PET - CT, clinical examinations (such as blood tests or urine tests), or physical examinations with high reliability and low false - positive rates. The diagnosis of the pathogen - related disorder can be made by a physician certified based on the results of the aforementioned or other well - established clinical examinations. In some cases, the result of the first screening assay is diagnosed by a well - established clinical diagnostic test that the subject does not have the disorder, so there is no result of medical treatment of the subject for the pathogen - related disorder. Based on the evaluated risk, in some cases, the method is related to the pathogen associated with the subject

[0026] Based on the evaluated risk, in some cases, the method is related to the pathogen ​Including determining the frequency of a screening assay of a subject. The screening assay frequency may be correlated with risk, and the interval between two screening assays, e.g., the screening assay described herein and a subsequent follow-up screening assay, may be inversely correlated with risk. In certain cases, the method includes receiving data from a first screening assay performed at a first time point. The first screening assay may include determining characteristics of cell-free nucleic acid molecules from a pathogen in a biological sample from the subject. For example, the first screening assay may include obtaining a biological sample from the subject, the biological sample including cell-free nucleic acid molecules, e.g., cell-free DNA, from the subject and potentially from the pathogen. The first screening assay may also include determining characteristics of the cell-free nucleic acid molecules from a pathogen in the biological sample. Non-limiting characteristics of cell-free nucleic acid molecules from a pathogen that can be used in the methods and systems provided herein include amount (e.g., copy number or percentage), methylation status, fragment size, mutation pattern, and relative abundance compared to cell-free nucleic acid molecules from the subject in the biological sample. As described herein, a time point for an examination or assay performed on a subject or a biological sample from the subject refers to the time when the subject undergoes the examination or when the biological sample is obtained from the subject, rather than the time when the actual assay is performed on the biological sample.

[0027] ​​​​​​​​​​​​​In one case, the method provided herein is performed at a first time point that includes (a) determining the characteristics of cell-free nucleic acid molecules from a pathogen in a biological sample of the subject receiving data from a first assay that includes determining the characteristics of cell-free nucleic acid molecules from a pathogen in a biological sample of the subject, wherein the characteristics of the cell-free nucleic acid molecules from the pathogen include quantity (e.g., copy number or percentage), methylation status, mutation pattern, fragment size, or relative abundance compared to cell-free nucleic acid molecules from the subject in the biological sample, and wherein the characteristics indicate the risk that the subject will develop a pathogen-related disorder ; and (b) based on the characteristics, determining a second time point at which a second assay is performed to screen for the pathogen-related disorder in the subject, wherein the interval between the first time point and the second time point is inversely correlated with the risk . In one case, one or more characteristics of the cell-free nucleic acid molecules in the biological sample of the subject described herein enable a non-invasive approach for assessing the status of the pathogen-related disorder (e.g., cancer ) in the subject or the risk that the subject will develop the pathogen-related disorder in the future. Without wishing to be bound by a particular theory, there are at least two possible scenarios underlying the relationship between one or more characteristics of the cell-free nucleic acid molecules that can be used in the methods and systems and the risk that the subject will develop the pathogen-related disorder . In one possible scenario, a diseased tissue affected by a pathogen-related disorder, e.g., a pathogen-related tumor, has a first screening (e.g.,

[0028] In one case, one or more characteristics of the cell-free nucleic acid molecules in the biological sample of the subject described herein enable a non-invasive approach for assessing the status of the pathogen-related disorder (e.g., cancer ) in the subject or the risk that the subject will develop the pathogen-related disorder in the future. Without wishing to be bound by a particular theory, there are at least two possible scenarios underlying the relationship between one or more characteristics of the cell-free nucleic acid molecules that can be used in the methods and systems and the risk that the subject will develop the pathogen-related disorder . In one possible scenario, a diseased tissue affected by a pathogen-related disorder, e.g., a pathogen-related tumor, has a first screening (e.g., a pathogen-related disorder, e.g., a pathogen-related tumor, has a first screening (e.g., ​​​​​​may already exist at the time of the first screening assay. However, the lesion tissue, for example, the size of the tumor is too small, and the false positive rate of detecting the pathogen-related disorder by other classical health diagnosis approaches, such as endoscopy and magnetic resonance imaging (MRI), etc. is less than 10%, 5%, 2%, 1%, 0.5%, 0.1%, or 0.05%, and may not be picked up by the approach. With the onset of the disorder, for example, with the growth of the lesion tissue, such as the size of the tumor, more advanced lesion tissue, such as enlarged tissue (for example, enlarged tumor), can be detected in subsequent screening (the second screening assay). Another possible scenario is as follows: the nucleic acid molecule of the pathogen, such as , EBV DNA, can be released by cells in a preliminary pathological condition such as pre-cancerous cells, and later, these cells may potentially develop into diseased cells such as cancer cells. Regardless of the exact scenario underlying the relationship, the subject matter described herein can be used to stratify subjects for the risk of having clinically detectable NPC subsequently. In certain cases, the actual time intervals used in the specific screening programs described in the specification are based on medical and economic considerations (such as the cost of screening), the preferences of the subjects (for example, more frequent screening intervals may be more disruptive to the lifestyle of a particular subject), and other clinical parameters (such as an individual's genetic profile (such as HLA status (Bei et al. Nat Genet. 2010; 42:599-603; Hildesheim et al.) ).

[0029] In certain cases, the actual time intervals used in the specific screening programs described in the specification are based on medical and economic considerations (such as the cost of screening), the preferences of the subjects (for example, more frequent screening intervals may be more disruptive to the lifestyle of a particular subject), and other clinical parameters (such as an individual's genetic profile (such as HLA status (Bei et al. Nat Genet. 2010; 42:599-603; Hildesheim et al.) and other clinical parameters (such as an individual's genetic profile (such as HLA status (Bei et al. Nat Genet. 2010; 42:599-603; Hildesheim et al. J Natl Cancer Inst. 2002; 94:1780-9.), family history of NPC, diet history, ethnic origin (e.g., Cantonese))).

[0030] In certain cases, the methods provided herein involve: receiving data from a first assay that includes determining a characteristic of a cell-free nucleic acid molecule from a pathogen in a biological sample of the subject, wherein the characteristic of the cell-free nucleic acid molecule from the pathogen includes an amount (e.g., copy number or percentage), methylation status, mutation pattern, fragment size,[ coordinates of the fragment ends, sequence motif of the fragment ends, or relative abundance compared to cell-free nucleic acid molecules from the subject in the biological sample; and generating a report indicating the risk that the subject will develop a pathogen-related disorder based on the characteristic of the cell-free nucleic acid molecule from the pathogen and one or more factors of: the age of the subject, the smoking habit of the subject, the family history of pathogen-related disorders of the subject, the genotype factors of the subject, or the diet history of the subject.

[0031] In an aspect, provided herein are methods and systems for analyzing nucleic acid molecules in a biological sample from a subject. Examples of the methods and systems include analyzing the mutation pattern of nucleic acid molecules from a pathogen in the biological sample. In certain cases, the nucleic acid molecules from the pathogen in the biological sample include cell-free nucleic acid molecules. The mutation pattern analysis includes comparing the sequence of the nucleic acid molecules in the biological sample identified as originating from the pathogen to one or more reference genomes of the pathogen, and subsequently the biological s ample ​​​​​​​​​​​Determining the nucleotide mutation pattern in the nucleic acid molecule from the pathogen in the sample; may include.

[0032] In certain cases, the methods and systems provided herein are based on the mutation pattern of the nucleic acid molecule from the pathogen in the biological sample to determine the condition or risk of pathogen-related disorders in the subject. For example, genetic mutations in the EBV genome detected in plasma can be used to predict the risk of future NPC development. Although it has been previously reported that the strains of EBV present in EBV-related tumors and control samples may be different (Palser et al. J Virol 2015; 89:5222-37), the tumor and control samples in this study were collected from geographically different locations. Therefore, considering the geographical variation of EBV variants, it is difficult to conclude whether the variants identified in tumor samples are geographically related or disease-related. EBV-related tumors and the EBV strains present in control samples may be different, as previously reported (Palser et al. J Virol 2015; 89:5222-37). However, the tumor and control samples in this study were collected from geographically different locations. Therefore, considering the geographical variation of EBV variants, it is difficult to conclude whether the variants identified in tumor samples are geographically related or disease-related. It has been previously reported that the strains of EBV present in EBV-related tumors and control samples may be different (Palser et al. J Virol 2015; 89:5222-37). However, the tumor and control samples in this study were collected from geographically different locations. Therefore, considering the geographical variation of EBV variants, it is difficult to conclude whether the variants identified in tumor samples are geographically related or disease-related. Although it has been previously reported that the strains of EBV present in EBV-related tumors and control samples may be different (Palser et al. J Virol 2015; 89:5222-37), the tumor and control samples in this study were collected from geographically different locations. Therefore, considering the geographical variation of EBV variants, it is difficult to conclude whether the variants identified in tumor samples are geographically related or disease-related. In certain cases, the mutation pattern analysis described herein includes a genome-wide comparison between the nucleic acid molecule from the pathogen in the biological sample and one or more reference genomes of the pathogen. The genome-wide comparison may include sequence alignment across the entire genome of the pathogen and subsequent clustering analysis of nucleotide mutation patterns. In certain cases, the genome-wide comparison includes the analysis of nucleotide variants at multiple sites across the reference genome of the pathogen. These sites can potentially include all sites across the entire genome of the pathogen. Alternatively, the reference genome of the pathogen

[0033] In certain cases, the mutation pattern analysis described herein includes a genome-wide comparison between the nucleic acid molecule from the pathogen in the biological sample and one or more reference genomes of the pathogen. The genome-wide comparison may include sequence alignment across the entire genome of the pathogen and subsequent clustering analysis of nucleotide mutation patterns. In certain cases, the genome-wide comparison includes the analysis of nucleotide variants at multiple sites across the reference genome of the pathogen. These sites can potentially include all sites across the entire genome of the pathogen. Alternatively, the reference genome of the pathogen The genome-wide comparison may include sequence alignment across the entire genome of the pathogen and subsequent clustering analysis of nucleotide mutation patterns. In certain cases, the genome-wide comparison includes the analysis of nucleotide variants at multiple sites across the reference genome of the pathogen. These sites can potentially include all sites across the entire genome of the pathogen. Alternatively, the reference genome of the pathogen The genome-wide comparison may include sequence alignment across the entire genome of the pathogen and subsequent clustering analysis of nucleotide mutation patterns. The genome-wide comparison may include sequence alignment across the entire genome of the pathogen and subsequent clustering analysis of nucleotide mutation patterns. These sites or variant sites across the reference genome are typically those where nucleotide variants can be found at at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 11 00, at least 1200, at least 1300, at least 1400, at least 15 00, at least 1600, at least 1700, at least 1800, at least 19 00, at least 2000, at least 3000, at least 4000, or at least 5000 sites. The nucleotide variants described herein can include single nucleotide variants (SNVs). The variant sites used in the variant pattern analysis provided herein can include typical SNVs identified in the genome of the pathogen. In certain cases, the variant sites can include insertions, deletions, and fusions.

[0034] The genome-wide variant pattern analysis provided herein may be superior to the analysis of individual single nucleotide polymorphisms (SNPs). In one exemplary case, SNPs on a fixed number of sites can be associated with a particular strain or subtype of the pathogen that may lead to the pathology of the subject, but the risk assessment based on the analysis of these individual SNPs may be limited to a particular strain or subtype of the pathogen, and if there are strains or subtypes that cause other diseases of the pathogen, the risk assessment may be at risk of being limited to a particular strain or subtype of the pathogen, and if there are strains or subtypes that cause other diseases of the pathogen, the risk assessment may be limited to the risk of may be insufficient to provide an accurate assessment. In another exemplary case, the genome-wide mutation pattern analysis provided in this specification is beneficial when the pathogen nucleic acid molecules in the biological sample are insufficient, for example, when analyzing cell-free nucleic acid molecules in a biological sample such as plasma. The available pathogen nucleic acid molecules in the biological sample may not cover a substantial amount of the pathogen genome. As a result, a genome-wide mutation pattern analysis that includes a large number of mutation sites across the entire genome of the pathogen can provide a more comprehensive and comparative reading of the genotype characteristics of the cell-free nucleic acid molecules from the pathogen in the biological sample. On the other hand, an analysis that includes a fixed number of individual polymorphisms is limited to relatively small regions or several small regions of the genome, and thus can provide a relatively limited reading of the genotype characteristics of the cell-free nucleic acid molecules from the pathogen in the biological sample. In some cases, the mutation pattern analysis provided in this specification includes block-based pattern analysis, which involves separating the reference genome of the pathogen into a plurality of bins and analyzing the sequence reads associated with each of the plurality of bins. In some cases, the method includes determining a similarity index for each of the plurality of bins relative to the disorder-related reference genome of the pathogen. The similarity index may be correlated with the proportion of mutation sites within each bin where at least one of the sequence reads mapped to the reference genome of the pathogen has the same nucleotide variant as the disorder-related reference genome of the pathogen. In some cases, the disorder-related reference genome of the pathogen is a plurality of the disorders of the pathogen of the cell-free nucleic acid molecules from the pathogen in the biological sample. can provide a relatively limited reading of the genotype characteristics.

[0035] In some cases, the mutation pattern analysis provided in this specification includes block-based pattern analysis, which involves separating the reference genome of the pathogen into a plurality of bins and analyzing the sequence reads associated with each of the plurality of bins. In some cases, the method includes determining a similarity index for each of the plurality of bins relative to the disorder-related reference genome of the pathogen. The similarity index may be correlated with the proportion of mutation sites within each bin where at least one of the sequence reads mapped to the reference genome of the pathogen has the same nucleotide variant as the disorder-related reference genome of the pathogen. In some cases, the disorder-related reference genome of the pathogen is a plurality of the disorders of the pathogen of the cell-free nucleic acid molecules from the pathogen in the biological sample. is mapped to the reference genome of the pathogen, and at least one of the sequence reads has the same nucleotide variant as the disorder-related reference genome of the pathogen. The similarity index may be correlated with the proportion of mutation sites within each bin. In some cases, the disorder-related reference genome of the pathogen is a plurality of the disorders of the pathogen . In some cases, the disorder-related reference genome of the pathogen is a plurality of the disorders of the pathogen comprising a plurality of disorder-related reference genomes of the pathogen, the method comprising determining a respective similarity index for each of the plurality of bins with respect to each of the plurality of disorder-related reference genomes of the pathogen; and determining a bin score for each of the plurality of bins based on the proportion of the plurality of disorder-related reference genomes in which the respective similarity index within each respective bin exceeds a cut-off value.

[0036] Assay of cell-free nucleic acid molecules The screening assay of the cell-free nucleic acid molecules from the biological sample of the subject can be any suitable nucleic acid assay. For example, sequencing methods can be employed to analyze quantity (e.g., copy number or percentage), methylation status, fragment size, or relative abundance of the cell-free nucleic acid molecules. Alternatively or additionally, amplification or hybridization-based methods, such as various polymerase chain reaction (PCR) methods or microarray-based approaches, can also be used. In certain cases, for example, to analyze the methylation status of the nucleic acid molecules, immunoprecipitation methods are used.

[0037] In certain examples of the present disclosure, the screening assay for detecting the cell-free pathogen nucleic acid molecules, such as cell-free EBV DNA, comprises two or more tests performed at various time points, and the detectability of the cell-free pathogen nucleic acid molecules over the plurality of tests can indicate the risk that the subject develops the pathogen-related disorder. For example, the assay comprises a two-step assay or three, four, five, six, seven, eight, nine, ten, or even more tests. It can include an assay regimen. Some tests can be run at the same time, while other tests can be run at different times, or all tests can be run at different times.

[0038] The timing or screening frequency of different screening assays can be determined by the methods and systems provided herein. The interval between the first screening assay and the second screening assay can be at least about 2 months, 4 months, 6 months, 8 months, 10 months, or 12 months. In some cases, the interval is at least about 12 months. The interval between the first screening assay and the second screening assay can be about 1 year, 1.5 years, 2 years, 2.5 years, 3 years, 3.5 years, 4 years, 4.5 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, or more. As long as the subject is ordinarily diagnosed as not having the pathogen-related disorder by a well-established clinical diagnostic method (e.g., not having a clinically detectable pathogen-related disorder), the first screening assay can yield a positive result indicating the presence of the pathogen-related disorder, but the interval can be longer. The methods and systems provided herein can, for example, predict the risk that the subject will develop the pathogen-related disorder in the future, such as within 6 months, 12 months, 2 years, 3 years, 5 years, or 10 years. Based on the

[0039] evaluated risk, an appropriate follow-up time point can be determined. ​And / or can be optimized to improve specificity. In certain embodiments, the sample can be obtained immediately prior to performing the assay (e.g., a first sample can be obtained prior to performing a first assay, and a second sample can be obtained after performing the first assay and prior to performing a second assay). In certain embodiments, the sample can be obtained and stored for a period of time (e.g., several hours, days, or weeks) prior to performing the assay. In certain embodiments, the assay can be performed on the sample 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 3 months, 4 months, 5 months, 6 months, within 1 year, or more than 1 year after obtaining the sample from the subject. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. and a second sample can be obtained after performing the first assay and prior to performing a second assay). In certain embodiments, the sample can be obtained and stored for a period of time (e.g., several hours, days, or weeks) prior to performing the assay. In certain embodiments, the assay can be performed on the sample 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 3 months, 4 months, 5 months, 6 months, within 1 year, or more than 1 year after obtaining the sample from the subject. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. The time between performing an assay (e.g., a first assay or a second assay) and determining whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay.

[0040] After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay. After performing an assay (e.g., a first assay or a second assay), the time to determine whether the sample contains a marker or set of markers indicative of a disorder such as a tumor can vary. In certain instances, the time can be optimized to improve the sensitivity and / or specificity of the assay or method. In certain embodiments, determining whether the sample contains a marker or set of markers indicative of a tumor can occur up to 0.1 hour, 0.5 hour, 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or within 1 month after performing the assay.

[0041] The sequencing analysis of the biological samples described herein can be performed for the analysis of one or more characteristics of cell-free nucleic acid molecules from a pathogen. The methods provided herein are for The sequencing analysis of the biological samples described herein can be performed for the analysis of one or more characteristics of cell-free nucleic acid molecules from a pathogen. The methods provided herein are for Nucleic acid molecules from biological samples, such as cell-free nucleic acid molecules, cellular nucleic acid molecules, or both may include sequencing. In one example, the methods provided herein analyze the results of sequencing, such as sequencing reads, from nucleic acid molecules from biological samples . The methods and systems provided herein may or may not include active steps of sequencing. The methods and systems may include or provide means for receiving and processing sequencing data from a sequencer . The methods and systems may also include or provide means for giving commands to adjust the parameters of the sequencing process to the sequencer, such as commands based on the analysis of sequencing results . . . .

[0042] Commercially available sequencing apparatuses, such as Illumina sequencing platforms and 454 / Roche platforms, can be used in the methods provided by this disclosure. Nucleic acid sequencing can be performed using any method known in the art. For example, sequencing may include next-generation sequencing. In one example, nucleic acid sequencing may include chain termination sequencing, hybridization sequencing, Illumina sequencing (e.g., using reversible terminator dyes), ion torrent semiconductor sequencing, mass spectrophotometry sequencing, ultra-parallel signature sequencing . . . . . . . ​Determination (MPSS) (massively parallel signature sequencing), Maxam-Gilbert sequencing, nanopore sequencing, polony sequencing, pyrosequencing, shotgun sequencing, single molecule real-time (SMRT) sequencing, SOLiD sequencing (hybridization using four fluorescently labeled dibase probes), universal sequencing, or any combination thereof can be used. One sequencing method used in the methods provided herein can include, for example, paired end sequencing using Illumina's "Paired End Module" with its genome analyzer. After the genome analyzer completes the first sequencing read, the paired end module can direct a second round of re-synthesis and cluster generation of the original template. By using paired end reads in the methods provided herein, sequence information can be obtained from both ends of a nucleic acid molecule and both ends can be mapped to a reference genome, such as the genome of a pathogen or the genome of a host organism. After mapping both ends, a pathogen integration profile can be determined according to some embodiments of the methods provided herein.

[0043] One sequencing method used in the methods provided herein can include, for example, paired end sequencing using Illumina's "Paired End Module" with its genome analyzer. After the genome analyzer completes the first sequencing read, the paired end module can direct a second round of re-synthesis and cluster generation of the original template. By using paired end reads in the methods provided herein, sequence information can be obtained from both ends of a nucleic acid molecule and both ends can be mapped to a reference genome, such as the genome of a pathogen or the genome of a host organism. After mapping both ends, a pathogen integration profile can be determined according to some embodiments of the methods provided herein.

[0044] ​​​​​​​​​​​​​During paired-end sequencing, the sequence read from the first end of the nucleic acid molecule is at least at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 1 15, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 165, at least 170, at least 175, or at least 1 80 consecutive nucleotides. The sequence read from the first end of the nucleic acid molecule is at most 24, at most 28, at most 32, at most 38, at most 42, at most 48, at most 52, at most 58, at most 62, at most 68, at most 72, at most 78, at most 82, at most 88, at most 92, at most 98, at most 102, at most 108, at most 122, at most 128, at most 132, at most 138, at most 142, at most 148, at most 152 at most 158, at most 162, at most 168, at most 172, or at most 180 consecutive nucleotides. The sequence read from the first end of the nucleic acid molecule is about 2 0, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 7 0, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 10 5, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 15 5, about 160, about 165, about 170, about 175, or about 180 consecutive nucleotides It may include. The sequence read from the second end of the nucleic acid molecule is at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 105, at least 12 0, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 155, at least 160, at least 1 65, at least 170, at least 175, or at least 180 consecutive nucleotides. The sequence read from the second end of the nucleic acid molecule is at most 24, at most 28, at most 32, at most 38, at most 42, at most 48, at most 52, at most 58, at most 62, at most 68, at most 72, at most 78, at most 82, at most 88, at most 9 2, at most 98, at most 102, at most 108, at most 122, at most 128, at most 1 32, at most 138, at most 142, at most 148, at most 152, at most 158, at most 162, at most 168, at most 172, or at most 180 consecutive nucleotides. The sequence read from the second end of the nucleic acid molecule is about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 105, about 120, about 12 5, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 16 5. It may contain 170, 175, or about 180 consecutive nucleotides. In some cases the sequence read from the first end of the nucleic acid molecule may contain at least 75 consecutive nu cleotides. In some cases, the sequence read from the second end of the nucleic acid molecule may contain at least 75 consecutive nucleotides. The sequence reads from the first end and the second end of the nucleic acid molecule may be of the same length or different lengths. The sequence reads from multiple nucleic acid molecules from a biological sample may be of the same length or different lengths .

[0045] Sequencing in the methods provided herein can be carried out at various sequencing depths . Sequencing depth can refer to the number of times a locus is covered by sequence reads aligned to the locus. A locus can be as small as a nucleotide, as large as an arm of a chromosome, or as large as the entire genome. The sequencing depth in the methods provided herein can be 50-fold, 100-fold, etc., where the number before "x" refers to the number of times the locus is covered by the sequence reads . Sequencing depth can also be applied to multiple loci or the entire genome, in which case x refers to the average number of times each locus or haploid genome or the entire genome is sequenced . In some cases, ultra-deep sequencing is carried out by the methods described herein, which means it can be carried out at a sequencing depth of at least 100-fold .

[0046] During the sequencing process (e.g., sequencing depth), a specific nucleotide within the nucleic acid is read ​The number of times taken or the average number of times can be several times larger than the length of the nucleic acid whose sequence is determined. In one example, the sequence depth is sufficiently larger than the length of the nucleic acid (e.g., at least 5-fold ) and the sequencing can be called "deep sequencing". In one example, the sequence depth can be on average at least about 5-fold, at least about 1 0-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 5 0-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 9 0-fold, at least about 100-fold larger than the length of the nucleic acid to be sequenced. In one example, the sample can be concentrated for a particular analyte (e.g., a nucleic acid fragment, or a cancer-specific nucleic acid fragment).

[0047] The sequence reads (or sequencing reads) generated in the methods provided herein refer to a string of nucleotides sequenced from any part or all of a nucleic acid molecule. For example, a sequence read can refer to a short string of nucleotides complementary to a nucleic acid fragment (e.g., 20 - 150), a string of nucleotides complementary to the ends of a nucleic acid fragment, or a string of nucleotides complementary to all the nucleic acid fragments present in a biological sample. Sequence reads can be obtained by various methods, such as sequencing techniques.

[0048] Quantity / detectability One of the characteristics of the cell-free nucleic acid molecules that can be used in the methods and systems described above is the amount (e.g., copy number or percentage) of cell-free nucleic acid molecules from a pathogen. The present Certain aspects of the disclosure relate to stratifying the risk that a subject will develop a pathogen-related disorder based on an assessment of the amount of cell-free nucleic acid molecules from a pathogen in a biological sample from the subject (e.g., copy number or percentage).

[0049] The copy number of nucleic acid molecules in a biological sample can be related to the detectability of the nucleic acid molecules. In particular considering a given assay, the detectability of a nucleic acid template can correlate with the copy number of the template molecules, e.g., copy numbers below the lower limit of detection of the assay may be undetectable, but those above the lower limit of detection of the assay can be defined as "detectable." For example, quantitative polymerase chain reaction (qPCR) assays can typically have a limit of detection where the signal of the template molecules cannot be distinguished from background noise. Thus, in some cases, the methods and systems provided herein directly depend on the detectability of cell-free nucleic acid molecules in a biological sample, which can be related to their copy number in the biological sample. In some cases, the copy number of cell-free nucleic acid molecules in a biological sample is directly measured. In other cases, the copy number is implicitly measured or inferred through the detection of the cell-free nucleic acid molecules themselves.

[0050] Detection assays such as polymerase chain reaction (PCR) or quantitative PCR (qPCR) can be performed to assess the presence or copy number of cell-free nucleic acid molecules from a pathogen in a biological sample. The probes can be designed to target genomic regions specific to the pathogen, e.g., genomic DNA sequences specific to Epstein-Barr virus (EBV), human papillomavirus (HPV), or hepatitis B virus (HBV).

[0051] Examples and embodiments are provided herein, for example, copy number and NPC Additional techniques and embodiments related to can be found in PCT filed on November 30, 2011, AU / 2011 / 001562, which is hereby incorporated by reference in its entirety into this specification. NPC may be closely related to EBV infection. In southern China, E BV genomes can be found in tumor tissues of almost all NPC patients. Plasma EBV DNA derived from NPC tissues has been developed as a tumor marker for NPC (Lo et al. Cancer Res 1999;59:1188-1191). In particular, real-time qPCR assays can be used for plasma EBV DNA analysis targeting the BamHI-W fragment of the EBV genome . Each EBV genome contains approximately 6 to 12 repeats of the BamHI-W fragment, and each NP C tumor cell may contain approximately 50 EBV genomes (Longnecker et al. Fields Virology , 5th Edition Chapter 61 “Epstein-Barr virus”; Tierney et al. JVirol. 2011; 85 :12362-12375). In other words, each NPC tumor cell may contain approximately 300 to 600 (e.g., approximately 500) copies of the PCR target. This large number of targets per tumor cell can explain why plasma EBV DNA is a highly sensitive marker in the detection of early NPC. NPC cells can deposit fragments of EBV DNA into the bloodstream of the subject. This tumor marker is used for monitoring NPC . This tumor marker is for monitoring NPC The ring (Lo et al. Cancer Res 1999;59:5452-5455) and prognostic diagnosis (Lo et al. Cance r Res 2000;60:6878-6881) are useful.

[0052] The qPCR assay can also be used in a manner similar to that described herein for EBV to measure the amount of HPV, HBV, or any other viral DNA A in a sample. Such an analysis is particularly useful for screening cervical cancer (CC), head and neck squamous cell carcinoma (HNSCC) , cirrhosis, or hepatocellular carcinoma (HCC). In one example, the qPCR assay targets a region (e.g., 200 nucleotides in length) within the polymorphic L1 region of the HPV genome. More specifically, contemplated herein is the use of qPCR primers that selectively hybridize to sequences encoding one or more hypervariable surface loops in the L1 region.

[0053] Alternatively, cell-free nucleic acid molecules from a pathogen can be detected and quantified using sequencing techniques. For example, cfDNA fragments can be sequenced and aligned to an HPV reference genome for quantification. Or, in another example, sequence reads of cfDNA fragments can be aligned to an EBV or HBV reference genome for quantification.

[0054] The detectability or copy number of cell-free nucleic acid molecules from a pathogen measured by the assays provided herein may indicate the risk that a subject will develop a pathogen-related disorder. In one example, the higher the copy number of cell-free nucleic acid molecules from a pathogen, the more likely the subject is to develop a pathogen-related disorder. There is a tendency for the risk of onset to increase. In one example, the detectability of cell-free nucleic acid molecules from a pathogen through one or more assays over one specific time point or multiple time points indicates the risk that the subject will develop a pathogen-related disorder. If the cell-free nuclear molecules from the pathogen in the biological sample from the subject can be detected by the assays provided herein as compared with the case where they cannot be detected, the subject has a higher tendency to be at risk of developing a pathogen-related disorder. The multi-step detection assay can be performed at the timing as described above. In one example of the present disclosure, a two-step assay is performed to detect cell-free pathogen nucleic acid molecules in a biological sample. In one case, the first test of the two-step assay is performed, and then, depending on the assay result at the first time point, the second test of the two-step assay is performed or not. As an example, if the result of the first test is positive, for example, if cell-free pathogen nucleic acid molecules are detected in the first biological sample, the second test of the two-step detection assay can be performed; if the result of the first test is negative, the second test may not be performed. In other cases, the second test is performed regardless of the first test. In one example, if positive results are obtained in both tests of the two-step detection assay, it is referred to as continuously positive, and if positive results are obtained in only the first or the second test, it is referred to as temporarily positive. In one exemplary example, a "positive" assay result indicates that the subject has a high risk of developing a pathogen-related disorder, such as EBV-related

[0055] NPC, as compared with a "negative" assay result, while a "continuously positive" assay result is In one case, the first test of the two-step assay is performed, and then, depending on the assay result at the first time point, the second test of the two-step assay is performed or not. As an example, if the result of the first test is positive, for example, if cell-free pathogen nucleic acid molecules are detected in the first biological sample, the second test of the two-step detection assay can be performed; if the result of the first test is negative, the second test may not be performed. In other cases, the second test is performed regardless of the first test. In one example, if positive results are obtained in both tests of the two-step detection assay, it is referred to as continuously positive, and if positive results are obtained in only the first or the second test, for example, if cell-free pathogen nucleic acid molecules are detected in the first biological sample, the second test of the two-step detection assay can be performed; if the result of the first test is negative, the second test may not be performed. In other cases, the second test is performed regardless of the first test. In one example, if positive results are obtained in both tests of the two-step detection assay, it is referred to as continuously positive, and if positive results are obtained in only the first or the second test, the second test may not be performed. In other cases, the second test is performed regardless of the first test. In one example, if positive results are obtained in both tests of the two-step detection assay, it is referred to as continuously positive, and if positive results are obtained in only the first or the second test, the second test is performed. In one example, if positive results are obtained in both tests of the two-step detection assay, it is referred to as continuously positive, and if positive results are obtained in only the first or the second test, it is referred to as temporarily positive. In one exemplary example, a "positive" assay result indicates that the subject has a high risk of developing a pathogen-related disorder, such as EBV-related NPC, as compared with a "negative" assay result, while a "continuously positive" assay result is indicates that the subject has a high risk of developing a pathogen-related disorder, such as EBV-related NPC, as compared with a "negative" assay result, while a "continuously positive" assay result is Indicates a higher risk compared to an assay result of "temporarily positive". In one exemplary example when the result is temporarily positive, if a persistent positive result is obtained from a two - step detection assay performed at a first time point, a longer interval can be set between the first and second time points. For example, in EBV - related NPC screening when a continuously positive result is obtained from the first two - step detection assay, it may be recommended to perform a follow - up screening assay within about one year from the first detection assay . In contrast, when a temporarily positive result is obtained from the first two - step detection assay, a follow - up screening assay can be performed within about two years from the first detection assay. If a negative result is obtained, a follow - up screening assay can be spaced four years or more apart. In some cases, the selection of the interval where a preceding positive result indicating high risk will be overridden by a subsequent result indicating low risk can be overridden . For example, if a persistent positive result is obtained in the first year, and then for the next four years, regardless of the results obtained from the performed follow - up assays, the subject will be followed up annually for the next four years . An exemplary example is given in Figure 2 and described in more detail in Example 2 . Similar to the detection assay, risk assessment based on other characteristics of cell - free nucleic acid molecules from pathogens can also follow this exemplary or similar screening regimen . The second test of the assay can be performed hours, days, or weeks after the first assay . In one example, the second assay is performed immediately after the first assay . . . . An exemplary example is given in Figure 2 and described in more detail in Example 2 . Similar to the detection assay, risk assessment based on other characteristics of cell - free nucleic acid molecules from pathogens can also follow this exemplary or similar screening regimen .

[0056] . The second test of the assay can be performed hours, days, or weeks after the first assay . In one example, the second assay is performed immediately after the first assay It is possible. In other cases, the second assay can be performed 1 day, 2 days, 3 days , 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 7 weeks, 8 weeks later, within 3 months, 4 months, 5 months, 6 months, within 1 year, or more than 1 year after the first assay. In a specific example, the second assay can be performed within 2 weeks from the first sample. Generally, the second test of the assay can be used to improve the specificity by which a pathogen-related disorder, such as a tumor, can be detected in a patient. The time between performing the first test and performing the second test can be determined experimentally. In one embodiment, the method can include two or more tests , and both tests use the same sample (e.g., a single sample is obtained from a subject, such as a patient, before performing the first assay and is stored for the period until the second assay is performed). For example, two blood tubes can be obtained from a subject simultaneously. The first tube can be used for the first test. The second tube can be used only if the result of the first test from the subject is positive. The sample can be stored using any method known to those skilled in the art (e.g., at cryogenic temperatures). This storage can be beneficial in certain situations, such as when a subject can receive a positive test result (e.g., the first assay indicates cancer), and the patient cannot wait until the second assay is performed, but rather when seeking a second opinion. The sample can be stored using any method known to those skilled in the art (e.g., at cryogenic temperatures). This storage can be beneficial in certain situations, such as when a subject can receive a positive test result (e.g., the first assay indicates cancer), and the patient cannot wait until the second assay is performed, but rather when seeking a second opinion. (e.g., the first assay indicates cancer), and the patient cannot wait until the second assay is performed, but rather when seeking a second opinion. (e.g., the first assay indicates cancer), and the patient cannot wait until the second assay is performed, but rather when seeking a second opinion.

[0057] Methylation status Certain aspects of the present disclosure relate to stratifying the risk that a subject will develop a pathogen-related disorder based on an assessment of the methylation status of cell-free nucleic acid molecules from a pathogen in a biological sample derived from the subject. ​

[0058] Methylation of cell pathogen nucleic acid molecules can identify samples from patients with pathogen-related disorders (e.g., EBV-related NPC or HPV-related cervical cancer) and subjects without such disorders (e.g., non-NPC subjects ). For example, the methylation status of plasma EBV DNA A can be different from the methylation status of plasma EBV DNA detected in non-NP C subjects, as shown in U.S. Patent Application No. 16 / 046,795, which is incorporated herein by reference in its entirety. When analyzed by bisulfite sequencing, there may be regions with different methylation between plasma DNA from NPC patients and non-NPC subjects with detectable EBV DNA. As a result, analysis of the methylation status in these different methylation regions can identify NPC and non-NPC subjects. As described in this specification, the NPC-related EBV DNA methylation status can also predict the risk of NPC development and can be used to adjust the interval of NPC screening. For example, subjects with a certain NPC-related EBV DNA methylation pattern can be screened more frequently than subjects without an NPC-related EBV DNA methylation pattern. In some cases, instead of bisulfite sequencing, for example, Pacific Bi osciences (Kelleher et al. Methods Mol Biol. 2018;1681:127-137; Powers et al. B MC Genomics. 2013;14:675) and Oxford Nanopore (Simpson et al. Nat Methods. 20 ). As described in this specification, the NPC-related EBV DNA methylation status can also predict the risk of NPC development and can be used to adjust the interval of NPC screening. For example, subjects with a certain NPC-related EBV DNA methylation pattern can be screened more frequently than subjects without an NPC-related EBV DNA methylation pattern. In some cases, instead of bisulfite sequencing, for example, Pacific Bi osciences (Kelleher et al. Methods Mol Biol. 2018;1681:127-137; Powers et al. B MC Genomics. 2013;14:675) and Oxford Nanopore (Simpson et al. Nat Methods. 20 ). osciences (Kelleher et al. Methods Mol Biol. 2018;1681:127-137; Powers et al. B MC Genomics. 2013;14:675) and Oxford Nanopore (Simpson et al. Nat Methods. 20 single-molecule sequencing systems such as (17;14:407-10), and methylation before sequencing Using restriction enzyme treatment, another type of methylation recognition sequencing can be performed. Furthermore In another case, a molecular approach that recognizes methylation and is not based on sequencing, such as methylation-specific PCR (Herman et al. Proc Natl Acad Sci U S A. 1996;93:9821-6), detection systems based on methylation-sensitive enzymes (e.g., restriction enzymes) and bisulfite conversion, followed by mass spectrometry (van den Boom et al. Methods Mol Biol. 2009;507:207-27 ; Nygren et al. Clin Chem. 2010; 56:1627-35), and approaches based on differential precipitation of DNA molecules based on methylation status or methyl-binding proteins (Zhang et al. Nat Commun. 2013; 4:1517) can be used (Shen et al . Nature. 2018; 563:579-83; Zhou et al. PLoS One. 2018; 13:e0201586). In one case, the methylation pattern of cell-free pathogen nucleic acid molecules, such as plasma EBV DNA, can be used for the detection of pathogen-related disorders, such as pathogen-related cancers such as NPC, or the prediction of future risks of clinically detectable disorders. As described above, one approach is to treat nucleic acid molecules with bisulfite to convert unmethylated cytosine to uracil. Methylated cytosine is bisulfite resistant

[0059] In one case, the methylation pattern of cell-free pathogen nucleic acid molecules, such as plasma EBV DNA, can be used for the detection of pathogen-related disorders, such as pathogen-related cancers such as NPC, or the prediction of future risks of clinically detectable disorders. As described above, one approach is to treat nucleic acid molecules with bisulfite to convert unmethylated cytosine to uracil. Methylated cytosine is bisulfite resistant approach is to treat nucleic acid molecules with bisulfite to convert unmethylated cytosine to uracil. Methylated cytosine is bisulfite resistant It remains cytosine without being changed by the treatment. After the bisulfite treatment of the nucleic acid molecule, subsequent assays, such as sequencing, can be employed to detect the methylation status of the nucleic acid molecule in a biological sample. In one example, the difference in the methylation level of plasma EBV DNA is determined using methylation-sensitive restriction enzyme analysis. One non-limiting example of a methylation-sensitive restriction enzyme is HpaII, which can cleave molecules with an unmethylated "CCGG" motif but does not modify molecules without "CCGG" or with methylated "CCGG". Alternatively or additionally, other methylation-sensitive restriction enzymes can be used. In one example, since the methylation level of plasma EBV DNA in non-cancer subjects is low, the plasma EBV DNA in non-cancer subjects may be more sensitive to cleavage by methylation-sensitive restriction enzymes. The susceptibility to enzymatic digestion can be determined by, for example, but not limited to, massively parallel sequencing, gel electrophoresis, capillary electrophoresis, polymerase chain reaction (PCR), and real-time PCR.

[0060] When using sequencing such as massively parallel sequencing to analyze the degree of digestion by methylation-sensitive restriction enzymes, regardless of the presence or absence of enzymatic digestion, the size distribution of cell-free nucleic acid molecules of a pathogen, such as plasma EBV DNA, can be used to reflect the degree of digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. One non-limiting example of a methylation-sensitive restriction enzyme is HpaII, which can cleave molecules with an unmethylated "CCGG" motif but does not modify molecules without "CCGG" or with methylated "CCGG". In one example, since the methylation level of plasma EBV DNA in non-cancer subjects is low, the plasma EBV DNA in non-cancer subjects may be more sensitive to cleavage by methylation-sensitive restriction enzymes. "CCGG" motif. Alternatively or additionally, other methylation-sensitive restriction enzymes can be used. In one example, since the methylation level of plasma EBV DNA in non-cancer subjects is low, the plasma EBV DNA in non-cancer subjects may be more sensitive to cleavage by methylation-sensitive restriction enzymes. The susceptibility to enzymatic digestion can be determined by, for example, but not limited to, massively parallel sequencing, gel electrophoresis, capillary electrophoresis, polymerase chain reaction (PCR), and real-time PCR. When using sequencing such as massively parallel sequencing to analyze the degree of digestion by methylation-sensitive restriction enzymes, regardless of the presence or absence of enzymatic digestion, the size distribution of cell-free nucleic acid molecules of a pathogen, such as plasma EBV DNA, can be used to reflect the degree of digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. When using sequencing such as massively parallel sequencing to analyze the degree of digestion by methylation-sensitive restriction enzymes, regardless of the presence or absence of enzymatic digestion, the size distribution of cell-free nucleic acid molecules of a pathogen, such as plasma EBV DNA, can be used to reflect the degree of digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion.

[0061] When using sequencing such as massively parallel sequencing to analyze the degree of digestion by methylation-sensitive restriction enzymes, regardless of the presence or absence of enzymatic digestion, the size distribution of cell-free nucleic acid molecules of a pathogen, such as plasma EBV DNA, can be used to reflect the degree of digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. As shown in FIGS. 12 and 13, a shift of the size distribution curve to the left can indicate a shortening of the size distribution of plasma EBV DNA. The greater the shift of the curve to the left, the higher the degree of enzymatic digestion. and means that the methylation level of the DNA becomes lower.

[0062] The methylation state of the cell-free pathogen nucleic acid molecules described herein is with respect to individual methylation sites the methylation density, the distribution of methylated / unmethylated sites across adjacent regions on the pathogen's genome, within one or more specific regions on the pathogen's genome, or the pattern or level of methylation of individual methylated sites across the entire genome of the pathogen, and may include non-CpG methylation. In one case, the methylation state includes the methylation level (or methylation density) of individual identified methylation sites, which can be identified, for example, between samples from patients with a pathogen-related disorder (such as EBV-related NPC or HPV-related cervical cancer) and subjects without the disorder ( for example, non-NPC subjects). The methylation density can refer to the fraction of nucleic acid molecules methylated at a given methylation site with respect to the total number of nucleic acid molecules of interest containing such a methylation site. For example, the methylation density of the first methylation site in liver tissue can refer to the fraction of liver DNA molecules methylated at the first site with respect to the total liver DNA molecules. In one case, the methylation state includes the coherence (such as pattern or haplotype) of the methylated / unmethylated state between individual methylation sites.

[0063] In one case, the screening assays described herein (such as the first assay or the second assay) can be any available technique, such as, but not limited to, the determination of methylation recognition sequences, the implementation of methylation-sensitive amplification or methylation-sensitive precipitation It may include determining the methylation state of acid molecules. Examples and embodiments are provided herein although, for example, additional techniques and embodiments related to determining methylation states can be found in PCT AU / 2013 / 001088 filed on September 20, 2013 and are hereby incorporated herein by reference in their entirety.

[0064] Fragment size Certain aspects of the present disclosure relate to stratifying the risk that a subject will develop a pathogen-related disorder based on an assessment of the fragment size of cell-free nucleic acid molecules from a pathogen in a biological sample derived from the subject.

[0065] The fragment size distribution and / or relative abundance of cell-free pathogen nucleic acid molecules can distinguish samples from patients with pathogen-related disorders (e.g., EBV-related NPC or HPV-related cervical cancer) and subjects without the disorder (e.g., non-NPC subjects). For example, the size distribution of plasma EBV DNA molecules and the ratio of EBV genome and human genome circulating DNA molecule mapping can be useful for distinguishing NPC patients from non-NPC subjects with detectable plasma EBV DNA, as shown using massively parallel sequencing in La m et al. Proc Natl Acad Sci U S A. 2018;115:E5115-E5124, which is hereby incorporated herein by reference in its entirety. According to certain examples of the present disclosure, the NPC-related size distribution and relative abundance of circulating DNA mapping to EBV and the human genome can also be useful for predicting the risk of developing clinically detectable NPC in the future. In one implementation, having these NPC-related characteristics for plasma DNA sequencing but being detectable ​Subjects without capable NPCs have detectable plasma EBV DNA but do not have these NPC -related features can be followed more frequently than subjects without such features. Using this sequencing-based analysis rather than the two-step assay described above to stratify the risk of NPC One potential practical advantage is that collection of another blood sample from the patient can be omitted.

[0066] In some cases, the assay (e.g., the first assay or the second assay) may include performing an assay, such as a next-generation sequencing assay, to analyze nucleic acid fragment size, e.g., the fragment size of plasma EBV DNA. In some cases sequencing is used to evaluate the size of cell-free viral nucleic acids in the sample. For example, the size of each sequenced plasma DNA molecule can be derived from the start and end coordinates of the sequence, and the coordinates can be determined by mapping (aligning) the sequence reads to the viral genome. In various examples, the start and end coordinates of the DNA molecule can be determined from two paired-end reads or a single read covering both ends, as can be achieved in single-molecule sequencing. In some cases, amplification or hybridization -based methods can also be used for fragment size analysis. For example, probes can be designed to target genomic regions of various lengths, and amplification (e.g., PCR or qPCR) or hybridization signals can indicate the number of cell-free nucleic acid fragments in the target genomic region while having a length equal to or greater than that of the target region. Thus, the distribution of fragment sizes can be estimated. -ization-based methods can also be used for fragment size analysis. For example, probes can be designed to target genomic regions of various lengths, and amplification (e.g., PCR or qPCR) or hybridization signals can indicate the number of cell-free nucleic acid fragments in the target genomic region while having a length equal to or greater than that of the target region. Thus, the distribution of fragment sizes can be estimated. It can be done. Methods for assay and analysis of fragment size can include those described in U.S. Patent Publication No. US20180208999A1, which are incorporated herein by reference in their entirety .

[0067] The fragment size distribution can be displayed as a histogram with the size of the nucleic acid fragment on the x-axis and. Determine the number of nucleic acid fragments at each size (e.g., within a resolution of 1 bp) and can be plotted on the y-axis, for example, as the raw number or percentage of frequency . The resolution of the size can be greater than 1 bp (e.g., a resolution of 2, 3, 4, or 5 bp ). The following analysis of the size distribution (also called the size profile) shows that viral DNA fragments in the cell-free mixture from NPC subjects are statistically longer than those from subjects without observable pathology . In one exemplary example, in the fragment size distribution curve obtained from plasma EBV DNA analysis, there may be a peak of 166 bp (nucleosome pattern) characteristic of the plasma EBV DNA size profile in NPC patients, while plasma EBV DNA from non-cancer subjects does not show a typical nucleosome pattern.

[0068] In some cases, to assess the risk, the relative abundance of cell-free nucleic acid molecules from a pathogen compared to cell-free nucleic acid molecules from a subject is calculated. In some cases, the relative abundance is analyzed from the perspective of the size ratio. In various examples, the size ratio of pathogen fragments to cell-free fragments from a subject refers to the quantitative ratio between cell-free nucleic acid fragments from a pathogen and cell-free nucleic acid fragments from a subject. For example, for example, 80 to 110 base pairs The size ratio of EBV DNA fragments between them can be as follows:

Number

[0069] In various cases, a cut-off value or threshold is set for evaluation. For example, there may be a size threshold for determining the size ratio between a pathogen fragment and a subject's autosomal fragment. Alternatively, in some cases, a size threshold is set such that several fragments having a size below or above the threshold are considered indicative of the subject's risk of developing a pathogen-related disorder. It should be understood that the size threshold can be any value. The size threshold can be at least about 10 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. For example, the size threshold can be 150 bp. In another example, the size threshold can be 180 bp. In some embodiments, upper and lower size thresholds can be used (e.g., a range of values). In some embodiments, upper and lower size thresholds are used to identify nucleic acid fragments having a length between the upper and lower cut-off values. between a pathogen fragment and a subject's autosomal fragment. There may be a size threshold. Alternatively, in some cases, several fragments having a size below or above the threshold are considered indicative of the subject's risk of developing a pathogen-related disorder. such that several fragments having a size below or above the threshold are considered indicative of the subject's risk of developing a pathogen-related disorder. It should be understood that the size threshold can be any value. The size threshold can be at least about 10 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. It should be understood that the size threshold can be any value. The size threshold can be at least about 10 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. p, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. p, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. 5 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. p, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. 0 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. It should be understood that the size threshold can be any value. The size threshold can be at least about 10 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. For example, the size threshold can be 150 bp. In another example, the size threshold can be 180 bp. It should be understood that the size threshold can be any value. The size threshold can be at least about 10 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 105 bp, 110 bp, 115 bp, 120 bp, 125 bp, 130 bp, 135 bp, 140 bp, 145 bp, 150 bp, 155 bp, 160 bp, 165 bp, 170 bp, 175 bp, 180 bp, 185 bp, 190 bp, 195 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, or 250 bp or more. For example, the size threshold can be 150 bp. In another example, the size threshold can be 180 bp. In some embodiments, upper and lower size thresholds can be used (e.g., a range of values). In some embodiments, upper and lower size thresholds are used to identify nucleic acid fragments having a length between the upper and lower cut-off values. It is possible to select. In certain embodiments, upper and lower cut - offs are used to select nucleic acid fragments having a length longer than the upper cut - off value and shorter than the lower size threshold It is possible to select. In certain cases, a cut - off value of the size ratio is used to determine whether there is a risk to a subject or the degree of risk that a subject will develop a pathogen - related disorder, such as NPC. For example, a subject with NPC has a lower size ratio within the size range of 80 - 110 bp compared to a subject who has obtained a false - positive result for plasma EBV DNA A. In certain cases, the cut - off value of the size ratio can be greater than about 0.1, about 0.5, about 1 , about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13 , about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 50, about 10 0, or greater than about 100. In certain cases, the cut - off value of the size index can be about or at least 10, about or at least 2, about or at least 1 , about or at least 0.5, about or at least 0.333, about or at least 0.25 , about or at least 0.2, about or at least 0.167, about or at least 0. 143, about or at least 0.125, about or at least 0.111, about or at least 0.1, about or at least 0.091, about or at least 0.083, about or at least 0.077, about or at least 0.071, about or at least 0.06 7, about or at least 0.063, about or at least 0.059, about or at least 0.056, about or at least 0.053, about or at least 0.05, about or at least 0.04, about or at least 0.02, about or at least 0.001, or less. In certain cases, the cut - off value of the size ratio can be greater than about 0.1, about 0.5, about 1 less. In certain cases, the cut - off value of the size ratio can be greater than about 0.1, about 0.5, about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 50, about 100, or greater than about 100. In certain cases, the cut - off value of the size index can be about or at least 10, about or at least 2, about or at least 1, about or at least 0.5, about or at least 0.333, about or at least 0.25, about or at least 0.2, about or at least 0.167, about or at least 0.143, about or at least 0.125, about or at least 0.111, about or at least 0.1, about or at least 0.091, about or at least 0.083, about or at least 0.077, about or at least 0.071, about or at least 0.067, about or at least 0.063, about or at least 0.059, about or at least 0.056, about or at least 0.053, about or at least 0.05, about or at least 0.04, about or at least 0.02, about or at least 0.001, or less. or can be less than about 0.001.

[0070] Various statistical values of the size distribution of the nucleic acid fragments can be determined. For example, representative values, modal values, median values, or average values of the size distribution can be used. Other statistical values, such as the cumulative frequency of a given size, or various ratios of the amounts of nucleic acid fragments of various sizes, can be used. The cumulative frequency can correspond to the proportion (e.g., percentage) of DNA fragments of a given size, less than a given size, or greater than a given size. The statistical values provide information regarding the distribution of the sizes of the nucleic acid fragments for comparison with one or more cutoffs for determining the level of pathology due to a pathogen. The cutoffs can be determined using cohorts of healthy subjects, subjects known to have one or more pathologies, subjects that are false positives for pathogen-related pathologies, and other subjects described herein. One of ordinary skill in the art will know how to determine such cutoffs based on the description herein. In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001. or can be less than about 0.001.

[0071] In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region). In one example, a first statistical value of the size of the pathogen fragment can be compared to a reference statistical value of the size from the human genome. For example, a separation value (e.g., a difference or ratio) can be determined between the first statistical value and a reference statistical value determined from another region of the pathogen reference genome or from a human nucleic acid. The separation value can also be determined from other values. For example, the reference value can be determined from the statistical values of multiple regions. The separation value is compared to a size threshold to obtain a size classification (e.g., whether the DNA fragment is shorter, longer, or the same as compared to a normal region).

[0072] In one example, a parameter (separation value) that can be defined as the difference in the ratio of short DN between the reference pathogen genome and the reference human genome can be calculated using the following formula: A parameter (separation value) that can be defined as the difference in the ratio of short DN between the reference pathogen genome and the reference human genome can be calculated:

Number

[0073] The size-based z-score can be calculated using the mean and SD values of the control subjects. can be.

Number

[0074] In one embodiment, a size-based z-score greater than 3 indicates an increase in the proportion of short fragments of the pathogen, while a size-based z-score less than 3 indicates a decrease in the proportion of short fragments of the pathogen. Other size thresholds can be used. Further details of the size-based approach are described in U.S. Pat. Nos. 8,620,593 and 8,741,811 and U.S. Patent Publication No. 2013 / 0237431, which are hereby incorporated by reference in their entireties. and U.S. Patent Publication No. 2013 / 0237431, which are hereby incorporated by reference in their entireties. In one embodiment, a size-based z-score greater than 3 indicates an increase in the proportion of short fragments of the pathogen, while a size-based z-score less than 3 indicates a decrease in the proportion of short fragments of the pathogen. Other size thresholds can be used. Further details of the size-based approach are described in U.S. Pat. Nos. 8,620,593 and 8,741,811 and U.S. Patent Publication No. 2013 / 0237431, which are hereby incorporated by reference in their entireties. respectively, each of which is incorporated by reference in its entirety.

[0075] To determine the size of the nucleic acid fragment, at least some examples of the present disclosure are directed to any single-molecule analysis platform capable of analyzing the chromosomal origin and the length of the molecule. To determine the size of the nucleic acid fragment, at least some examples of the present disclosure are directed to any single-molecule analysis platform capable of analyzing the chromosomal origin and the length of the molecule. It can function better. Examples of such platforms include electrophoresis, optical methods( For example, optical mapping and its variants, en.wikipedia.org / wiki / Optical_mapping# cite_note-Nanocoding-3, and Jo et al. Proc Natl Acad Sci USA. 2007; 104:2673- 2678), fluorescence-based methods, probe-based methods, digital PCR (microfluidic ics-based, or emulsion-based, for example, BEAMing (Dressman et al. Proc Natl Acad Sci USA. 2003; 100:8817-8822), RainDance(www.raindanc etech.com / technology / pcr-genomics-research.asp)), rolling circle amplification , mass spectrometry, melting analysis (or melting curve analysis), molecular sieves, etc. An example of mass spectrometry is that the longer the molecule, the larger the mass (an example of a size value).

[0076] In one example, nucleic acid molecules can be randomly sequenced using a paired-end sequencing protocol. The two reads at both ends can be mapped (aligned) to the reference genome and can be repeatedly masked (for example, when aligned to the human genome). The size of the DNA molecule can be determined from the distance between the genomic positions where the two reads map.

[0077] Analysis of mutation patterns Certain aspects of the present disclosure relate to stratifying the risk that a subject will develop a pathogen-related disorder based on an assessment of the mutation pattern of cell-free nucleic acid molecules from a pathogen in a biological sample derived from the subject. ​​​​It can be used to predict the future onset risk of pathogen-related disorders. The genetic mutations of the pathogen genome detected in biological samples can be used to predict the future onset risk of pathogen-related disorders.

[0078] The mutation pattern of the pathogen nucleic acid molecule may differ in the diseased tissues from patients with pathogen-related disorders (e.g., pathogen-related malignancies) compared with samples from subjects without pathogen-related disorders. It has been reported that the EBV strains present in EBV-related tumors and control samples (Palser et al. J Virol. 2015; 89:5 222-37) may be different. However, in these previous studies, the tumor and control samples were collected from geographically different locations. Considering the potential geographical variation of EBV variants, it is difficult to conclude whether the variants identified in tumor samples are geographically or disease-related. There have been attempts to identify NPC-related EBV variants through the analysis of NPC tumor samples. In one genome-wide association study (GWAS) (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049) analyzing NPC tumors and saliva samples from individuals without EBV-related diseases from the same geographical region, 29 polymorphisms (single nucleotide polymorphisms (SNPs) or indels) were identified with a false discovery rate less than the adjusted P of 0.05. These 29 NPC-related EBV variants were shown to be present in more than 90% of NPC cases, but only in 40-50% of control cases. Analysis of individual EBV polymorphisms for the development of NPC (Hui et al. Int J Cancer 2019, do There have been attempts to identify NPC-related EBV variants through the analysis of NPC tumor samples. NPC tumors and saliva samples from individuals without EBV-related diseases from the same geographical region were analyzed. In one genome-wide association study (GWAS) (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049), 29 polymorphisms (single nucleotide polymorphisms (SNPs) or indels) were identified with a false discovery rate less than the adjusted P of 0.05. These 29 polymorphisms were identified with a false discovery rate less than the adjusted P of 0.05. These 29 NPC-related EBV variants were shown to be present in more than 90% of NPC cases. These 29 NPC-related EBV variants were shown to be present in more than 90% of NPC cases, but only in 40-50% of control cases.

[0079] Analysis of individual EBV polymorphisms for the development of NPC (Hui et al. Int J Cancer 2019, do In contrast to (i.org / 10.1002 / ijc.32049; Feng et al. Chin J Cancer 2015; 34:61) Aspects of the present disclosure provide methods and systems for analyzing pathogen nucleic acid molecules for mutation patterns in a genome-wide manner. Further, rather than identifying disease-related EBV variants by analysis of tumor and cell line samples (Palser et al. J Virol. 2015; 89:5222-37 , Correia et al. J Virol. 2018; 92:e01132-18, Hui et al. Int J Cancer 2019, doi .org / 10.1002 / ijc.32049), aspects of the present disclosure analyze cell-free pathogen nucleic acid molecules such as blood (e.g., plasma or serum), nasal wash, nasal brush samples, or other body fluids obtained via non-invasive or minimally invasive procedures as compared to invasive biopsies of tumors, to provide methods and systems for analyzing pathogen mutation patterns. In one exemplary example the low abundance and fragmented nature of EBV DNA molecules in blood can present technical challenges for analysis. Non-invasively analyzing the mutation patterns of cell-free viral DNA molecules can enhance clinical applications including screening, predictive medicine, risk stratification, monitoring, and prognosis. In one example, the analysis can be used to identify subjects with various virus-related conditions, e.g., NPC patients and non-NPC subjects with detectable plasma EBV DNA in the context of screening. In another example, it can be used for disease or cancer risk prediction. Different approaches can be used to obtain mutation patterns. Non-limiting examples include

[0080] ​​​​​​The assay method may include ultra - parallel sequencing (MPS), Sanger sequencing (such as that used in Lorenzetti et al . J Clin Microbiol. 2012; 50:609 - 18), and microarray - based SNP analysis (such as that described in Wang et al. PNAS 2002; 99:15687 - 92 ), hybridization analysis, and mass spectrometry analysis. In one exemplary example, sequencing methods such as capture enrichment, target sequencing with MPS or Sanger Sequencing are used, and sequence reads are analyzed with reference to the pathogen's reference genome (e.g., EBV reference genome) nucleotide by nucleotide. The method may include obtaining sequence reads of cell - free nucleic acid molecules from a biological sample of a subject . The method may further include aligning the sequence reads to the reference genome of the pathogen . The method may further include analyzing the nucleotide mutation pattern across the pathogen's reference genome by analyzing the nucleotide mutations between the pathogen's reference genome and the sequence reads mapped to the pathogen's reference genome . The mutation patterns provided herein can characterize the nucleotide variants of the sequence reads mapped to the pathogen's reference genome at each of a plurality of mutation sites on the pathogen's reference genome . The plurality of mutation sites may be at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 40 0, at least 500, at least 600, at least 700, at least 800, at least across the pathogen's reference genome . The mutation patterns provided herein can characterize the nucleotide variants of the sequence reads mapped to the pathogen's reference genome at each of a plurality of mutation sites on the pathogen's reference genome . The plurality of mutation sites may be at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 40 0, at least 500, at least 600, at least 700, at least 800, at least 90, at least 100, at least 200, at least 300, at least 40 0, at least 500, at least 600, at least 700, at least 800, at least sites of at least 900, at least 1000, at least 1100, at least 1200 can be included. In some cases, the plurality of variant sites span at least 1000 sites across the reference genome of the pathogen. In some cases, the plurality of variant sites span at least 1100 sites across the reference genome of the pathogen. In some cases, the plurality of variant sites span at least 600 sites across the reference genome of the pathogen. In some cases the plurality of variant sites span at least 660 sites across the reference genome of the pathogen. In some cases, the plurality of variant sites are at least 30, 40, 50 selected from genomic sites as described in Table 6 associated with the EBV reference genome (AJ507799.2), 100, 150, 200, 250, 300, 350, 400, 450, 500, 550 or 600 sites. In some cases, the plurality of variant sites include genomic sites as described in Table 6 associated with the EBV reference genome (AJ507799.2).

[0081] In some cases, the mutation pattern of the cell-free nucleic acid molecule from the pathogen is characterized by nucleotide variants of the sequence reads mapped to the reference genome of the pathogen at each of a plurality of variant sites randomly selected from genomic sites as described in Table 6 associated with the EBV reference genome (AJ507799.2). In some cases, the methods provided herein include the step of randomly selecting a plurality of variant sites from genomic sites as described in Table 6 associated with the EBV reference genome (AJ507799.2). The method characterizes the nucleotide variants of the sequence reads mapped between the reference genome of the pathogen and the reference genome of the pathogen. In some cases, the methods provided herein include randomly selecting a plurality of variant sites from genomic sites as described in Table 6 associated with the EBV reference genome (AJ507799.2). The method includes the step of mapping between the reference genome of the pathogen and the sequence reads mapped to the reference genome of the pathogen By analyzing nucleotide mutations, it may further include analyzing nucleotide mutation patterns across a plurality of randomly selected mutation sites which may cover.

[0082] In certain cases, the mutation pattern of cell-free nucleic acid molecules from a pathogen is randomly selected from genomic sites as described in Table 6 related to the EBV reference genome ( AJ507799.2), at least 30, 40, 50, 100, 150, 200, 250, 300, 3 50, 400, 450, 500, 550, or 600 sites, including a plurality of mutation sites characterizing the nucleotide variants of the sequence reads mapped to the reference genome of the pathogen at each of them. In certain cases, the plurality of mutation sites consists of all sites where the sequence reads mapped to the reference genome of the pathogen have nucleotide variants different from the reference genome of the pathogen.

[0083] In certain cases, the wild-type pathogen genome is used as the reference genome. For example, the wild-type EBV genome (GenBank: AJ507799.2) can be used as the reference EBV genome. In other cases, other pathogen genomes are used as the reference genome. In yet another example, a plurality of pathogen genomes (e.g., EBV genomes) are used as references. In still another example, a consensus sequence is used as the reference. The consensus can be constructed by combining variants of different pathogen genome sequences, such as the consensus sequence of the EBV genome described in de Jesus et al. J Gen Virol. 2003; 84:1443-50.

[0084] ​

[0085] For example, for the analysis of copy number, methylation status, fragment size, relative abundance, or mutation pattern the alignment utilized in the methods and systems provided herein can be implemented by any suitable bioinformatics algorithm, program, toolkit, or package. For example, as an alignment tool for the application of the methods and systems provided herein, the Short Oligonucleotide Analysis Package (SOAP) can be used. Examples of short sequence read analysis tools that can be used in the methods and systems provided herein include Arioc, Barra CUDA, BBMap, BFAST, BigBWA, BLASTN, BLAT, Bowt ie, Bowtie2, BWA, BWA-PSSM, CASHX, Cloudburst 、CUDA-EC、CUSHAW、CUSHAW2、CUSHAW2-GPU、CUSH AW3, drFAST, ELAND, ERNE, GASSST, GEM, Generalic e MAP, Geneious Assembler, GensearchNGS, GMA P and GSNAP, GNUMAP, HIVE-hexagon, Isaac, LAST 、MAQ、mrFAST、mrsFAST、MOM、MOSAIK、MPscan、No voalign&NovoalignCS, NextGENe, NextGenMap, Omixon Variant Toolkit, PALMapper, Partek F low, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG Investigator, Segemehl 、MAQ、mrFAST、mrsFAST、MOM、MOSAIK、MPscan、No voalign&NovoalignCS, NextGENe, NextGenMap, Omixon Variant Toolkit, PALMapper, Partek Flow, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG Investigator, Segemehl , SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, S OAP3, SOAP3-dp, SOCS, SparkBWA, SSAHA, SSAHA2 , Stampy, STORM, Subread, Subjunc, Taipan, UGE NE, VelociMapper, XpressAlign, and ZOOM are included .

[0086] Some consecutive nucleotides within a sequence read (a "sequence stretch") can be used to align to a reference genome and make calls regarding the alignment . For example, the alignment can be to a reference genome, such as a reference genome of a pathogen, or also at least 4, at least 6, at least 8, at least 10, at least 12, at least 14, at least 16, at least 1 8, at least 20, at least 22, at least 24, at least 25, at least 2 6, at least 28, at least 30, at least 32, at least 34, at least 3 5, at least 36, at least 38, at least 40, at least 42, at least 4 4, at least 45, at least 46, at least 48, at least 50, at least 5 2, at least 54, at least 55, at least 56, at least 58, at least 6 0, at least 62, at least 64, at least 65, at least 66, at least 6 7, at least 68, at least 69, at least 70, at least 71, at least 7 2, at least 73, at least 74, at least 75, at least 76, at least 7 8, at least 80, at least 82, at least 84, at least 85, at least 8 6. At least 88, at least 90, at least 92, at least 94, at least 9 5. At least 96, at least 98, at least 100, at least 102, at least 104, at least 106, at least 108, at least 110, at least 112 . At least 114, at least 116, at least 118, at least 120, at least 122, at least 124, at least 126, at least 128, at least 13 0, at least 132, at least 134, at least 136, at least 138, at least 140, at least 142, at least 145, at least 146, at least 1 48, or at least 150 consecutive nucleotides may be aligned. In some cases, the alignment as referred to herein is a reference genome, e.g., a reference genome of a pathogen, or up to 5, up to 7, up to 9, up to 11, up to 13, up to 15, up to 17, up to 19, up to 21, up to 23, up to 25, up to 27, up to 29, up to 31, up to 33 , up to 37, up to 39, up to 41, up to 43, up to 45, up to 47, up to 49, up to 51, up to 53, up to 55, up to 57, up to 59, up to 61, up to 63, up to 65, up to 67, up to 68, up to 69, up to 70, up to 71 , up to 72, up to 73, up to 74, up to 75, up to 76, up to 78, up to 80, up to 81, up to 83, up to 85, up to 87, up to 89, up to 91, up to 93, up to 95, up to 97, up to 99, up to 101, up to 103, up to 105, up to 107, up to 109, up to 111, up to 113, up to 115, Up to 117, up to 119, up to 121, up to 123, up to 125, up to 12 7, up to 129, up to 131, up to 133, up to 135, up to 137, up to 139, up to 141, up to 143, up to 145, up to 147, up to 149, or may include aligning up to 151 consecutive nucleotides. In one example , the alignment as referred to herein is about 20, about 22, about 24, about 2 5, about 26, about 28, about 30, about 32, about 34, about 35, about 36, about 38, about 40, about 4 2, about 44, about 45, about 46, about 48, about 50, about 52, about 54, about 55, about 56, about 5 8, about 60, about 62, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 7 1, about 72, about 73, about 74, about 75, about 76, about 78, about 80, about 82, about 84, about 8 5, about 86, about 88, about 90, about 92, about 94, about 95, about 96, about 98, about 100, about 102, about 104, about 106, about 108, about 110, about 112, about 114, about 116, about 118, about 120, about 122, about 124, about 126, about 128, about 130, about 132, about 118, about 120, about 122, about 124, about 126, about 128, about 130, about 132, about 134, about 136, about 138, about 140, about 142, about 145, about 146, about 148, about 150, about 152, about 154, about 155, about 156, about 158, about 160, about 162, about 164, about 165, about 166, about 168, about 170, about 172, about 174, about 175, about 176, about 178, about 180, about 185, about 190, about 195, or about 200 consecutive nucleotides may be aligned.

[0087] In certain cases, sequence stretches that have at least a certain region of the reference genome across the entire sequence read, e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 99%, or 100% identity or complementarity to the human reference genome, an alignment is called. In certain cases , sequence stretches that have at least 80% identity or complementarity to a certain region of the reference genome across the entire sequence read, e.g., to the human reference genome, an alignment is called. In certain cases, sequence stretches that are identical or complementary to a certain region of the reference genome, e.g., to the human reference genome, and have no more than 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2 bases of mismatch, or 1 base, or zero bases of mismatch, an alignment is called . In certain cases, sequence stretches that are identical or complementary to a certain region of the reference genome, e.g., to the human reference genome, and have no more than 2 bases of mismatch, an alignment is called. The maximum number or percentage of mismatches, or the minimum number or percentage of similarities, can vary as a selection criterion depending on the purpose and context of the methods and systems provided herein. In certain cases, by aligning sequence reads to the reference genome of a pathogen, it is possible to have no more than 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base of maximum mismatch. The mapped sequence reads and the reference genome of the pathogen

[0088] A mismatch with the genome can indicate nucleotide variations in the pathogen genome sequence present in a biological sample, or in other cases, can also indicate sequencing errors. Without wishing to be bound by a particular theory, the identification of two or more nucleotide variants at a given genomic site in one biological sample may be due to sequencing errors or heterogeneity in the diseased cells from which the cell-free pathogen nucleic acid molecules are derived. In one case, if one, two, or more than three nucleotide variants are identified in a given biological sample, the nucleotide variants at the genomic site are excluded from the analysis.

[0089] In an exemplary example, targeted sequencing with capture enrichment is used to analyze cell-free viral DNA molecules in the circulation of NPC subjects and non-NPC subjects with detectable plasma EBV DNA. The capture probes can be designed to cover the entire EBV genome. In other cases, only a part of the EBV genome can be analyzed, and the capture probes are designed to cover only a part of the EBV genome. In the same analysis, the target genomic regions of the human genome, including the capture probes, can also be targeted. For example, probes targeting human common single nucleotide polymorphism (SNP) sites and human leukocyte antigen (HLA) SNPs can be included. In one embodiment, more probes can be designed to hybridize to other viral genome sequences, such as the HPV or HBV genome.

[0090] In one case, the mutation pattern of the pathogen genome is analyzed by directly comparing the consensus reads mapped to the reference genome with the reference genome. These analyses are facilitated ​ Available bioinformatics tools include MEGA4, MEGA5, CLUSTAL W, Phylip, RAxML, BEAST, PhyML, TreeView, MAFF T, MrBayes, BIONJ, MLTreeMap, Newick Utiliti es, Phylo.io, Phylogeny.fr, REALPHY, SuperTre e, ThePhylOgenetic Web repeater. Cluster analysis or phylogenetic analysis compares sequence reads mapped to a pathogen reference genome with one or more pathogen genomes obtained from diseased tissue or healthy subjects, or shown to be able or unable to cause a pathogen-related disorder, or shown to be effective or ineffective in causing a pathogen-related disorder. In an exemplary example, the methods and systems provided herein include block-based mutation pattern analysis. Block-based mutation pattern analysis may include separating a pathogen reference genome into a plurality of bins ("blocks"). Sequence reads mapped to the pathogen reference genome are compared to the disorder-related pathogen genomes in each of the plurality of bins. In some cases, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different pathogen genomes are compared for block-based analysis, the disorder-related pathogen genomes, and optionally, the pathogen

[0091] ​​​​​​​​​are incapable of causing relevant disorders or are known to be ineffective or Includes pathogen genomes shown (pathogen genomes not related to disorders). Block-based In the analysis, sequences that map to the pathogen reference genome are identified within each of multiple bins. Shared nucleotides between the genomes of the genomes of the pathogens associated or unassociated with the disease A similarity index is calculated based on the shared nucleotide variants. The similarity index is , at least one of the sequence reads mapped to the pathogen reference genome is a disorder-associated or the proportion of mutation sites that have the same nucleotide variant as in the non-disorder-associated pathogen genome. It depends on the similarity index for each pathogen genome to which the sequence reads are compared. Thus, the bin scores are calculated based on the similarity level as reflected by, for example, a similarity index. In one example, the bin score can be calculated by dividing the number of similarity indices by the number of similarity indices that exceed a certain cutoff. Similarity indices can range, for example, from about 0.6, 0.7, 0.75, 0.8, 0. You can set a cutoff of 85, 0.9, or 0.95. Similarity indexes that exceed the cutoff The number can indicate how "similar" a sequence read is to the pathogen genome it is compared to. Based on the above analysis, the pattern analysis is performed using the calculated similarity index or bin score. In this case, the method can be performed on a larger scale across the pathogen genome or on a portion of the pathogen genome. Cluster or phylogenetic analyses similar to those used in this study have been used to identify pathogen-related NPCs, such as EBV-associated NPCs. It is possible to follow a block-based analysis to predict the risk of developing a chronic disorder.

[0092] Risk Score Certain aspects of the present disclosure include the step of detecting cell-free nucleic acid molecules from a pathogen in a biological sample from a subject. Stratification of the risk that a subject develops a pathogen-related disorder based on a combinatorial consideration of one or more characteristics. In one case, a risk score indicating the risk that a subject develops a pathogen-related disorder, such as EBV-related nasopharyngeal carcinoma, is generated.

[0093] In one case, the present disclosure relates to stratification of the risk that a subject develops a pathogen-related disorder based on a combinatorial consideration of one or more characteristics of cell-free nucleic acid molecules from a pathogen in a biological sample from the subject, and one or more factors including the age of the subject, the smoking habit of the subject, the family history of NPC of the subject, the genotype factors of the subject, the diet history, or the ethnicity of the subject. There is a positive correlation between the positive rate of detection of plasma EBV DNA in subjects without clinically detectable NPC and the age of the subject. The smoking habit of the subject can increase the risk of NPC development in the subject. Subjects with a family history of NPC can increase the risk of developing NPC themselves. As shown in Bei et al. Nat Genet. 2010; 42:599-603 and Hildesheim et al. J Natl Cancer Inst. 2002; 94:1780-9, each of which is incorporated herein by reference in its entirety, genotype factors such as HLA status are also correlated with the risk of NPC. Further, the diet history can be correlated with the risk of NPC. For example, subjects who consume a large amount of salted fish may have a relatively high risk of NPC. Certain ethnicities such as Cantonese may also be associated with a high risk of developing NPC.

[0094] In one case, the method and system stratify the risk that a subject develops a pathogen-related disorder. Further includes creating a report to be presented. Such a report may have numerical risk scores value or risk assessment by value or category. In some cases, the report includes screening frequency or recommendations regarding future time points of a tracking screening assay. The report can be provided to the subject, a medical institution or medical professional providing services to the subject, or any relevant third party such as an insurance company. The report can be reviewed, evaluated or edited by a certified physician before or after release of the report. In some cases, a certified physician provides additional comments on the risk assessment or contributes to the final risk assessment based on their medical opinion or independent examination.

[0095] In some cases, the present disclosure provides a method for stratifying the risk of developing a pathogen-related disorder such as a pathogen-related proliferative disorder such as EBV-related NPC by using a classifier. Such a classifier can take one or more of the factors described herein as data inputs and provide an output including a risk score, which can indicate the risk that a subject will develop a pathogen-related disorder. One or more factors feedable to the classifier can include one or more characteristics of the cell-free pathogen nucleic acid molecule, one or more characteristics of the cell-free nucleic acid molecule from the pathogen in a biological sample from the subject, and one or more factors of the subject's age, the subject's smoking habit, the subject's family history of NPC, the subject's genetic factors, diet history and the subject's ethnicity. The risk score as the output of the classifier can indicate whether the subject is currently suffering from or will develop in the future a pathogen-related disorder. In some cases, the risk score indicates that the subject is currently ill Indicates the likelihood of suffering from a pathogen-related disorder. In one case, the risk score indicates the likelihood that the subject will develop a pathogen-related disorder within a future period, such as, but not limited to, within 1 year, within 2 years, within 3 years, within 4 years, within 5 years, within 10 years, or within 15 years. In one case, the classifier provides an output including the recommended screening frequency of a follow-up screening assay or a future time point. Such output can be in the form of clinical recommendations, or provided in a report to a third party such as the subject, a medical institution, a medical professional, or a medical insurance company as described above. As described herein, the classifier can refer to any algorithm that implements classification . In the present disclosure, the classifier can be a classification model constructed based on any suitable algorithm for predicting the risk of future onset of a pathogen-related disorder. Suitable algorithms

[0096] can include machine learning algorithms, and, without limitation, support vector machines (SVMs), Naive Bayes, logistic regression, random forests , decision trees, gradient boosting trees, neural networks, deep learning, linear / kernel SVMs, linear / non-linear regression, linear discriminant analysis, and other mathematical / statistical models. In one case, the classifier is trained on a labeled dataset including a plurality of input-output pairs. For example, the dataset is generated from the analysis results of samples of a number of subjects diagnosed as having or not having NPC. In these examples, the dataset includes one or more characteristics of plasma EBV DNA from these subjects. In one case, the classifier is trained on a labeled dataset including a plurality of input-output pairs. For example, the dataset is generated from the analysis results of samples of a number of subjects diagnosed as having or not having NPC. In these examples, the dataset includes one or more characteristics of plasma EBV DNA from these subjects. In these examples, the dataset includes one or more Factors (e.g., mutation patterns, methylation status, detectability / copy number, or fragment size), age, family history, smoking habits, ethnicity, or diet history, and further may include an input including a corresponding output indicating whether the corresponding subject has NPC. In an exemplary example, the classifier can be trained with a labeled dataset including a large number of input-output pairs, such as at least 10, 20, 50, 100, 200, 500, 1000, 20 00, 5000, 10000, or 20000 pairs.

[0097] In one example, a classification model is provided to predict the risk of future NPC onset in subjects having detectable plasma EBV DNA A using analysis of mutation patterns. The classification model can be a classifier constructed as follows using a support vector machine (SVM) algorithm: When a training dataset including n samples is given: (M1, Y1), …, (Mn, Yn) where Yi indicates the NPC status of sample i. Yi is 1 in the case of a sample from an NPC patient or -1 in the case of a sample from a subject without NPC; M i is a p-dimensional vector including the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites (e.g., 29 mutation sites related to NPC as shown in Table 6 or 661 mutation sites related to NPC), or alternatively, Mi can be a series of block-based mutant similarity scores (e.g., non-overlapping windows of 500 bp) with respect to a reference EBV mutant present in a subject known to have NPC.

[0098] By obtaining a set of coefficients (W having a p-dimensional vector) that satisfy the following, the training separates non-NPC groups and NPC groups within the dataset as accurately as possible to identify the "hyperplane" that can be: Criterion 1: W·M i -b≧1 (for subjects in the NPC group) and Criterion 2: W·M i -b≦1 (for subjects in the non-NPC group) Here, W is a p-dimensional vector of coefficients that determine the hyperplane; M is a matrix (p x n order) having p variants ( or block-based similarity scores) and n samples; b is the intercept.

[0099] The two criteria (i.e., Criterion 1 and 2) can also be described as follows: Yi(W * Mi - b)≧1 (Criterion 3) Here, Yi is either -1 (non-NPC) or 1 (NPC).

[0100] The margin distance (D) between Criterion 1 and 2 is:

Number

[0101] D is maximized by minimizing according to Criterion 3 JPEG2025084804000006.jpg923.

[0102] Based on this principle, the parameters (W and b) of the classifier can be determined. Therefore , using the trained parameters (W and b), the trained classifier implemented can be used to calculate the NPC risk score of the test sample.

[0103] In one exemplary example, the NPC risk score is calculated as the weighted sum of EBV genotypes at a fixed set of SNV sites across the viral genome (as an explanatory variable in a binary logistic regression model). In this example, a set of NPC-related SNVs is identified by analyzing the differences in EBV SNV profiles from NPC samples and non-NPC samples within the training set. The association between each variant across the EBV genome and NPC cases can be analyzed using, for example, Fisher's exact test. Then, for example, a fixed set of significant SNVs can be obtained by controlling the false discovery rate (FDR) at 5%. The NPC risk score for a test sample can be determined by the EBV genotype for this specific set of significant SNVs identified from sequencing data from plasma DNA samples from known NPC and non-NPC subjects. In some cases, the concentration of plasma EBV DNA molecules can be low, resulting in incomplete coverage of the entire EBV genome by the sequenced EBV DNA reads. The score can be formulated to be determined by the genotype pattern across all SNV sites covered by the plasma EBV DNA reads (e.g., by the available genotype information). To derive the NPC risk score, a subset of significant SNV sites covered by the plasma EBV DNA reads within the sample is first identified, and then the genotypic weighting (effect size) at each site can be determined within the subset of significant SNV sites. A logistic regression model can be constructed to provide information on the effect sizes of the risk genotypes at each SNV site for NPC: ​

number

number

[0104] Biological samples Biological samples used in the methods provided herein may be of any type, including living or dead. The biological sample may include any tissue or material derived from a subject. It is possible. A biological sample may contain nucleic acids (e.g., DNA or RNA) or fragments thereof. The nucleic acid in the sample may be cell-free nucleic acid. The sample may be a liquid sample or a solid sample (e.g., a cell or tissue sample). Biological samples include blood, plasma, serum, urine, oral rinse fluid, nasal wash fluid, nasal brush sample, vaginal fluid, fluid from a hydrocele (e.g., the testis), vaginal wash fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, discharge fluid from the nipple, bodily fluids such as aspirated fluid from various parts of the body (e.g., the thyroid, breast), etc. Fecal samples can also be used. In various examples cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) may be concentrated Most of the DNA in the concentrated biological sample can be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free ). The biological sample can be processed to physically disrupt tissue or cell structures (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution that may further contain enzymes, buffers, salts, surfactants, etc. used to prepare the sample for analysis.

[0105] The methods and systems provided herein can be used to analyze nucleic acid molecules in biological samples. The nucleic acid molecules can be either cellular nucleic acid molecules, cell-free nucleic acid molecules, or both. The cell-free nucleic acid used by the methods provided herein is a biological sample It can be a nucleic acid molecule outside the cells in the sample. Cell-free nucleic acid molecules can be present in various body fluids such as blood, saliva, semen, and urine. Cell-free DNA molecules can be generated by cell death in various tissues that can be caused by a healthy state and / or disease, such as viral infection or tumor growth. Cell-free nucleic acid molecules can contain sequences generated as a result of pathogen integration events.

[0106] The cell-free nucleic acid molecules, such as cell-free DNA, used in the methods provided herein can be present in plasma, urine, saliva, or serum. Cell-free DNA can occur naturally in the form of short fragments. Fragmentation of cell-free DNA can refer to the process by which high molecular weight DNA (such as DNA in the cell nucleus) is cleaved, broken, or digested into short fragments when the cell-free DNA molecule is generated or released. The methods and systems provided herein can, in some cases, analyze cellular nucleic acid molecules, such as cellular DNA from tumor tissue, or cellular DNA from white blood cells when the patient has leukemia, lymphoma, or myeloma. Samples taken from tumor tissue can be the subject of assays and analysis according to some examples of the present disclosure.

[0107] Subject (target) The methods and systems provided herein can be used to analyze samples from a subject, such as an organism, such as a host organism. The subject can be any human patient, such as a cancer patient, a patient at risk of cancer, or a patient with a family or personal cancer history of cancer. ​​​​​​​is possible. In some cases, the subject is at a particular stage of cancer treatment. In some cases, the subject may have cancer or may be suspected of having cancer. In some cases, it is unknown whether the subject has cancer.

[0108] In some cases, depending on the results of the screening assays provided herein, the subject will or will not receive treatment for a pathogen-related disorder. In one example, a first screening assay indicates a positive result that the subject has a high risk of developing a pathogen-related disorder, but the subject is diagnosed by subsequent diagnostic tests not to have a pathogen-related disorder (e.g., EBV-related NPC). In this case, the subject will not receive medical treatment, such as but not limited to, therapeutic agents (e.g., chemotherapy), radiation therapy, surgery, or any combination thereof. In another example, the subject is screened as having a high risk of developing a pathogen-related disorder (e.g., HPV-related cervical cancer), and is further diagnosed as having the disorder. As a result, the subject may receive medical treatment for the disorder, such as but not limited to, surgery, chemotherapy, radiation therapy, targeted therapy, immunotherapy, or any combination thereof. surgery, chemotherapy, radiation therapy, targeted therapy, immunotherapy, or any combination thereof.

[0109] Pathogen-related disorders to which the methods and systems provided herein can be applied may include proliferative disorders, such as cancer. The disorder may be related to or caused by a pathogen such as a virus, bacterium, or fungus. Viruses that may be related to the disorders described herein include EBV, Kaposi's sarcoma-associated herpesvirus (KSHV), HPV (e.g., but not limited to, HPV16, 18, 31, 33, 34, 35, 39, 45, 51, 52, 56, 58, 59, 66, 68, 70, 73, 82, 83, 84, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326, 327, 328, 329, 330, 331, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 7​​​​​​ 56, 58, 59, 66, 68, and 70) (Burd et al. Clin Microbiol Rev 2003:1 6:1-17), Merkel cell polyomavirus (MCPV), HBV, HCV, and human T-lymphotropic virus-1 (HTLV1). The corresponding pathogen-related cancers can include Burkitt lymphoma, Hodgkin lymphoma, immunosuppression-related lymphoma, T-cell lymphoma, and NK-cell lymphoma; nasopharyngeal or gastric carcinomas that can be associated with EBV. The corresponding pathogen-related cancers can include primary effusion lymphoma or Kaposi sarcoma that can be associated with KSHV. The corresponding pathogen-related cancers can include cervical cancer, head and neck cancer, or anogenital carcinomas that can be associated with HPV. The corresponding pathogen-related cancers can include Merkel cell carcinoma that is associated with MCPV. The corresponding pathogen-related cancers can include HCC that can be associated with HBV or hepatitis C virus (HCV). The corresponding pathogen-related cancers can include adult T-cell leukemia / lymphoma that can be associated with HTLV1. The subject can have any type of cancer or tumor, or can have a risk of developing any type of cancer or tumor. In one example, the subject can have nasopharyngeal cancer or nasal cavity cancer. In another example, the subject can have oropharyngeal cancer or oral cavity cancer. Non-limiting examples of cancers include, but are not limited to, adrenal cancer, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, blood cancer, bone cancer, brain tumor, breast cancer, bronchial cancer, cardiovascular cancer, cervical cancer, colon cancer, colorectal cancer, digestive system cancer, endocrine system cancer, endometrial cancer, esophageal cancer, eye cancer, gallbladder cancer, digestive tumor, hepatocellular carcinoma, kidney cancer, hematopoietic malignancy, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, mesothelioma, muscle cancer,

[0110] The subject can have any type of cancer or tumor, or can have a risk of developing any type of cancer or tumor. In one example, the subject can have nasopharyngeal cancer or nasal cavity cancer. In another example, the subject can have oropharyngeal cancer or oral cavity cancer. Non-limiting examples of cancers include, but are not limited to, adrenal cancer, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, blood cancer, bone cancer, brain tumor, breast cancer, bronchial cancer, cardiovascular cancer, cervical cancer, colon cancer, colorectal cancer, digestive system cancer, endocrine system cancer, endometrial cancer, esophageal cancer, eye cancer, gallbladder cancer, digestive tumor, hepatocellular carcinoma, kidney cancer, hematopoietic malignancy, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, mesothelioma, muscle cancer, endocrine system cancer, endometrial cancer, esophageal cancer, eye cancer, gallbladder cancer, digestive tumor, hepatocellular carcinoma, kidney cancer, hematopoietic malignancy, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, mesothelioma, muscle cancer, Myelodysplastic syndrome (MDS), myeloma, nasal cancer, nasopharyngeal cancer, nervous system cancer, lymphatic cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, pituitary cancer, prostate cancer, rectal cancer, renal pelvic cancer, genital system cancer, respiratory system cancer, sarcoma, salivary gland cancer, skeletal system cancer, skin cancer, small intestine cancer, gastric cancer, testicular cancer, laryngeal cancer, thymic cancer, thyroid cancer, tumor, urinary tract cancer, uterine cancer, vaginal cancer, or vulvar cancer. Lymphomas can be any type of lymphoma including B-cell lymphomas (e.g., diffuse large B-cell lymphoma, follicular lymphoma, small lymphocyte lymphoma, mantle cell lymphoma, marginal zone B-cell lymphoma, Burkitt lymphoma, lymphoplasmacytic lymphoma, hairy cell leukemia, or primary central nervous system lymphoma), or T-cell lymphomas (e.g., precursor T-lymphoblastic lymphoma or peripheral T-cell lymphoma). Leukemias can be any type of leukemia including acute leukemia or chronic leukemia. Types of leukemia include acute myeloid leukemia chronic myeloid leukemia, acute lymphoblastic leukemia, acute undifferentiated leukemia, or chronic lymphocytic leukemia . In some cases, a cancer patient does not have a specific type of cancer. In some cases, a patient may have a cancer other than breast cancer. Examples of cancers include not only cancers that do not cause solid tumors but also cancers that cause solid tumors . Further, any of the cancers referred to herein can be a primary cancer (e.g., a cancer named after the part of the body where it first began to grow), or a secondary or metastatic cancer (e.g., a cancer that originated from another part of the body).

[0111] A subject diagnosed by any of the methods described herein can be of any age .

[0112] ​​​It can be an adult, infant or child. In certain cases, the subject is 0, 1, 2, 3, 4, 5 , 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2 0, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33 , 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 6 0, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73 , 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 years old, or within that range (e.g., 2 - 20 years old, 20 - 40 years old or 40 - 90 years old) . Certain classes of patients who can benefit can be patients 40 years of age or older . Another certain class of patients who can benefit can be pediatric patients. Further, a subject diagnosed by any of the methods or compositions described herein can be male or female

[0113] In certain embodiments, the methods of the disclosure can detect a tumor or cancer in a subject, and the tumor or cancer has a geographic pattern of the disease. In one example, the subject can have an EBV - related cancer (e.g., nasopharyngeal cancer) that is prevalent in southern China ( e.g., Hong Kong SAR). In another example, the subject can have an HPV - related cancer (e.g., oropharyngeal cancer) that can be prevalent in the United States and Western Europe . In yet another example, the subject can have an HTLV - 1 that can be prevalent in groups of immigrants in southern Japan, the Caribbean, Central Africa , parts of South America, and parts of the southeastern United States may have a related cancer (e.g., adult T-cell leukemia / lymphoma).

[0114] None of the methods disclosed herein are to be performed on non-human subjects such as laboratory or farm animals, or cell samples derived from the organisms disclosed herein either. Non-limiting examples of non-human subjects include dogs, goats, guinea pigs, hamsters, mice, rabbits, non-human primates (e.g., gorillas, chimpanzees, orangutans, squirrel monkeys, or baboons ), rats, sheep, cows, or zebrafish.

[0115] Computer system Any of the methods disclosed herein can be performed and / or controlled by one or more computer systems In one example, any step of the methods disclosed herein can be performed and / or controlled in whole, individually, or sequentially by one or more computer systems Any computer system mentioned herein can utilize any suitable number of subsystems. In one embodiment, the computer system includes a single computer device, where the subsystems can be components of the computer device In other embodiments, the computer system can include multiple computer devices, each of which is a subsystem and has internal components. The computer system can include desktop and laptop computers, tablets, mobile phones and other mobile devices. The subsystems can be interconnected via a system bus. Additional subsystems can be provided

[0116] ​​​​​​​Including a lint remover, a keyboard, a storage device, and a monitor coupled to a display adapter. Peripheral devices and input / output (I / O) devices coupled to an I / O controller The vice can be connected to a computer system by any number of connections known in the art, such as input / output (I / O) ports (e.g., USB, FireWire (registered trademark) ). For example, an I / O port or an external interface (such as Ethernet, Wi-Fi i, etc.) can be used to connect the computer system to the Internet, a mouse input device, or a wide area network such as a scanner. Through the mutual connection via the system bus, the central processor can communicate with each subsystem, and the system memory or a storage device (such as a fixed disk like a hard drive or an optical disk) to control the execution of multiple instructions from and the information exchange between subsystems. The system memory and / or the storage device can embody a computer-readable medium. Another subsystem is a data collection device such as a camera, a microphone, an accelerometer, etc. Any data described in this specification can be output from one component to another component and can be output to the user.

[0117] The computer system can include a plurality of identical components or subsystems connected together, for example, by an external interface or an internal interface. In certain embodiments, the computer system, subsystem, or device can communicate via a network. In such a case, one computer can be regarded as a client. ​​and another computer can be considered a server, each of which is part of the same computer system. It may be part of the stem.

[0118] The present disclosure provides a method for stratifying risk of pathogen-associated disorders. FIG. 21 shows a computer-controlled system for detecting cell-free nucleic acid molecules or The sequence reads are analyzed to assess other factors associated with the risk of failure. and / or to generate a report indicating the risks described herein. 1 illustrates a computer system 1101 that is configured to run a program or other program. 1101 controls, for example, the sequencing of nucleic acid molecules from a biological sample, Various steps of bioinformatics analysis of sequencing data as described in the specification Integrates data collection, analysis and reporting of results, and data management. Implementing and / or controlling various aspects of the methods provided in this disclosure, including The computer system 1101 can be a user's electronic device or The electronic device may be a computer system located remotely to the mobile device. The device may be a portable electronic device.

[0119] The computer system 1101 includes a central processing unit (CPU, herein referred to as a “processor”). processor” and “computer processor” 1105, which A multi-core or multi-core processor, or multiple processors for parallel processing The computer system 1101 also includes a memory or memory location 1110 (e.g., , random access memory, read-only memory, flash memory), electronic storage units 1115 (e.g., a hard disk), and a communication device for communicating with one or more other systems. An interface 1120 (e.g., a network adapter), as well as a cache (cac he), other memory, data storage and / or peripherals such as electronic display adapters The memory 1110, the storage unit 1115, the interface 1120 and peripheral device 1125 are connected to a motherboard via a communication bus (solid line). The memory unit 1115 communicates with the CPU 1105 via a The computer system 1 may be a data storage unit (or a data repository). 101 is assisted by a communication interface 1120 to communicate with a computer network (" The network 1130 may be operatively coupled to In some cases, it is a telecommunications and / or data network. 30 may include one or more computer servers, which may be used in cloud computing. The network 1130 may, in some cases, With the aid of the computer system 1101, a peer-to-peer network 1101. It allows devices to act as either a client or a server.

[0120] The CPU 1105 is capable of executing a sequence of machine-readable instructions, which may be a program or The instructions may be stored in a memory location, such as memory 1110. The instructions may be directed to the CPU 1105, which then executes the instructions of this disclosure. The CPU 1105 can be programmed or otherwise configured to implement the method. Examples of operations performed by the CPU 1105 include fetch, decode, execute, and write-back.

[0121] The CPU 1105 can be part of a circuit such as an integrated circuit. One or more other components of the system 1101 can be included in the circuit. In one case, the circuit is an application specific integrated circuit (ASIC).

[0122] The memory unit 1115 can store files such as drivers, libraries, and saved programs. The memory unit 1115 can store user data such as user preferences and user programs. The computer system 1101 can include one or more additional data storage units external to the computer system 1101, such as being located on a remote server that communicates with the computer system 1101 via an intranet or the Internet in some cases.

[0123] The computer system 1101 can communicate with one or more remote computer systems via the network 1130. For example, the computer system 1101 can communicate with a user's remote computer system (e.g., a smartphone having an application installed that receives and displays the results of a sample analysis sent from the computer system 1101). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate, or a tablet. PCs (e.g., Apple (R) iPad, Samsung (R) Galaxy Tab), phones, smartphones (e.g., Apple (R) iPhone, A ndroid-enabled devices, Blackberry (R)), or personal de igital assistants. The user can access the computer sys tem 1101 via the network 1130.

[0124] The methods described herein can be implemented, for example, by machine (e.g., computer processor) executable code stored in any electronic memory location of a computer system 1101, such as the memory 1110 or the electronic storage unit 1115. The machine executable code or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 1105. In some cases, the code is retrieved from the storage unit 1115 and stored in the memory 1110 for immediate access by the processor 1105. In some situa tions, the electronic storage unit 1115 can be excluded beforehand, and the machine execut able instructions are stored in the memory 1110.

[0125] The code can be pre-compiled and configured beforehand or compiled at runtime for use on a machine having a processor adapted to execute the code. The code can be provided in a pro gramming language selected such that the code can be executed in a pre-compiled or com piled fashion.

[0126] Aspects of the systems and methods provided herein, such as the computer system 1101, can be implemented in programming. Various aspects of the technology are usually considered as a "product" or "manufactured article" in terms of the type of machine (or processor) executable code and / or associated data that is executed or implemented in a tangible form. Machine executable code can be stored in an electronic memory device such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. The "memory" type of media can include any or all of the tangible memories of a computer, processor, etc., or various semiconductor memories, tape drives, disk drives, etc. that can provide non-transitory storage at any time for software programming. All or part of the software may be communicated via the Internet or other various telecommunications networks. Such communication can, for example, enable the loading of software from one computer or processor to another, e.g., from a management server or host computer to an application server's computer platform. Thus, another type of media that can carry software elements includes wired and optical fixed telework, and includes light, electricity, and electromagnetic waves as used through the physical interface between local devices via various air links. Physical elements that carry such waves, such as wire or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, terms such as "readable medium" of a computer or machine are not limited to Refers to any medium involved in providing instructions to a sensor.

[0127] Thus, machine-readable media such as computer-executable code can take many forms and include, but are not limited to, tangible storage media, carrier wave media, or physical transmission media. Volatile memory media can be any memory device such as any computer that can be used to implement, for example, a database shown in the drawings and includes optical or magnetic disks such as those in a storage device of any computer. Volatile memory media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wires and optical fibers including wires that include buses within a computer system. Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communication. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, and other magnetic media, CD-ROMs, DVDs or DVD-ROMs, and other optical media, punched card paper tapes, other physical memory media with patterns of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, and other memory chips or cartridges, carrier waves that transmit data or instructions, cables or links that transmit such carrier waves, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions for execution to a processor. to a processor. and / or data. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions for execution to a processor. to a processor.

[0128] The computer system 1101 provides, for example, but not limited to, a graphical display of a pathogen integration profile, the genomic position of a pathogen integration breakpoint, a pathological classification (e.g., the type of disease or cancer and the cancer level), and the results of sample analysis such as treatment proposals or recommendations for preventive measures based on the pathological classification to a user interface (UI) 1140. The computer system 1101 includes or is communicable with an electronic display 1135 that includes the UI. Examples of the UI include, but are not limited to, a graphical user interface (GUI) and a web-based user interface. The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit unit 1105. The algorithms can control, for example, sequencing of nucleic acid molecules from a sample, direct collection of sequencing data, analysis of the sequencing data, performance of a block-based variant pattern analysis, assessment of risk, or generation of a report indicating risk. In some cases, as shown in FIG. 22, the sample 1202 can be obtained from a subject 1201 such as a human subject. The sample 1202 can be subjected to one or more of the methods described herein, such as performing an assay. In some cases, the assay can include hybridization, amplification, sequencing, labeling, epigenetically modifying bases, or any combination thereof. One or more results from the method are input to a processor 1204.

[0129] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit unit 1105. The algorithms can control, for example, sequencing of nucleic acid molecules from a sample, direct collection of sequencing data, analysis of the sequencing data, performance of a block-based variant pattern analysis, assessment of risk, or generation of a report indicating risk. In some cases, as shown in FIG. 22, the sample 1202 can be obtained from a subject 1201 such as a human subject.

[0130] The sample 1202 can be subjected to one or more of the methods described herein, such as performing an assay. In some cases, the assay can include hybridization, amplification, sequencing, labeling, epigenetically modifying bases, or any combination thereof. One or more results from the method are input to a processor 1204. In some cases, as shown in FIG. 22, the sample 1202 can be obtained from a subject 1201 such as a human subject. The sample 1202 can be subjected to one or more of the methods described herein, such as performing an assay. can be performed. One or more input parameters such as sample identification, subject identification, sample type, reference, or other information can be input to the processor 1204. One or more measurement metrics from the assay can be input to the processor 1204 such that the processor can generate results such as a pathological classification (e.g., diagnosis) or treatment recommendations. The processor can transmit the results, input parameters, measurement metrics, references, or any combination thereof to a visual display or graphical user interface such as display 1205. The processor 1204 can (i) transmit the results, input para meters, measurement metrics, or any combination thereof to the server 1207, (ii) receive the results, input parameters, measurement metrics, or any combination thereof from the server 1207, or (iii) perform any combination thereof. from the server 1207, and (iii) perform any combination thereof. from the server 1207, and (iii) perform any combination thereof. from the server 1207, and (iii) perform any combination thereof. is possible.

[0131] Aspects of the present disclosure can be implemented in the form of a control logic circuit using hardware (e.g., application specific integrated circuit or field programmable gate array) and / or using computer software with a modular or integrated manner generally programmable processor. As used herein, a processor includes a single core processor, a multi-core processor on the same integrated chip, or a plurality of processing device units on a single circuit board or networked. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will appreciate other ways and / or using combinations of hardware and hardware and software to implement the embodiments described herein. including or networked. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will appreciate other ways and / or using combinations of hardware and hardware and software to implement the embodiments described herein. or combinations of hardware and software to implement the embodiments described herein. Or they will know the method and recognize its true value.

[0132] Any of the software components or functions described in this application are implemented as software code executable by a processor using Java, C, C++, C#, Objective-C, Swift, or scripting languages such as Perl or Python using conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands for storage and / or transmission on a computer-readable medium. Suitable non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), magnetic media such as hard drives or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, and the like. The computer-readable medium can be any combination of such storage or transmission devices.

[0133] Such programs can also be encoded and transmitted using carrier signals adapted for transmission via various protocols including the Internet, wired, optical, and / or wireless networks. Thus, the computer-readable medium can be created using data signals encoded with such programs. Computer-readable media encoded with program code can be packaged together with compatible devices or provided separately from other devices (e.g., via Internet download). Any such computer-readable can be present on or in a product (e.g., a hard drive, a CD, or an entire computer system) and can be present on or in different computer products within a system or network. The computer system can include a monitor, a printer, or other suitable display for providing any of the results described herein to a user.

[0134] Any of the methods described herein can be implemented in whole or in part using a computer system that includes one or more processors configured to perform the steps. Accordingly, embodiments can be directed to a computer system configured to perform any of the steps of the methods described herein, where different components perform each step or each group of steps. Although presented as numbered steps, the steps of the methods herein can be performed simultaneously or in a different order. Further, some of these steps can be used in conjunction with some of the steps of other methods. Also, all or some of the steps can be optionally selected. Further, any step of any method can be implemented by a module, unit, circuit, or other approach for performing those steps.

[0135] Other embodiments The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0136] The methods described herein are in the context of the specific methodologies, protocols, subject matter, and It is to be understood that the techniques are not limited to array determination, and thus can vary. Also, the terminology used in this specification is for the purpose of describing only certain embodiments and is not intended to limit the scope of the methods and compositions described in this specification, but is to be understood as being limited only by the appended claims. Although some embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only by way of example. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the present disclosure. It is to be understood that various alternatives to the embodiments of the present disclosure described herein can be used when practicing the present disclosure. The scope of the following claims defines the scope of the disclosure, and it is intended that methods and structures within the scope of these claims, as well as their equivalents, be covered thereby.

[0137] With reference to exemplary applications for illustration, some aspects are described. Unless otherwise specified, any embodiment can be combined with any other embodiment. It is to be understood that numerous specific details, relationships, and methods are shown in order to provide a complete understanding of the features described in this specification. However, those skilled in the art will readily recognize that the features described in this specification can be practiced without one or more of the specific details, or in other ways. Since some acts can occur in a different order and / or concurrently with other acts or events, the features described in this specification are not limited by the illustrated order of acts or events. Further, aspects in accordance with the features described in this specification ​​To implement the methodology, not all of the illustrated acts or events are required .

Example

[0138] The following examples are provided to further illustrate some embodiments of the present disclosure but are not intended to limit the scope of the disclosure; by their exemplary nature, it will be understood that other procedures, methodologies or techniques known to those skilled in the art may alternatively be used .

[0139] [Example 1. NPC Screening in a Cohort of Over 20,000 Subjects over 4 Years This example describes a large-scale screening study conducted in a cohort of over 20,000 subjects over approximately 4 years. Figure 1 shows a diagram of the design of this study. In the first round of screening, using plasma EBV DNA analysis, over 20 ,000 men aged 40 - 62 years were screened for NPC. Subjects with detectable plasma EBV DNA were retested with a second blood sample set after a median of 4 weeks. This decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified ,000 male subjects aged 40 - 62 years were screened for NPC using plasma EBV DNA analysis. Subjects with detectable plasma EBV DNA were retested with a second blood sample set after a median of 4 weeks. This decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified decision was aimed at differentiating between NPC patients and patients without NPC but with detectable plasma EBV DNA. Previous studies have shown that the presence of plasma EBV DNA in subjects without NPC is usually a transient phenomenon. In two-thirds of these individuals, plasma EBV DNA becomes undetectable after a median of 2 weeks. Subjects with persistently positive plasma EBV DNA results were further investigated by nasal endoscopy and magnetic resonance imaging (MRI) of the nasopharynx to confirm or rule out the presence of NPC. Based on this decision, 34 cases of NPC were identified Based on this decision, 34 cases of NPC were identified​

[0140] Then, another round of NPC screening (round 2) was performed on the cohort. The second round of NPC screening was performed a median of 4 years after the first round of screening. In this round, subjects who tested positive were screened in the same way as in the first round of screening. If two consecutive tests over a four-week period are positive, Subjects with nasal stenosis will be further investigated with nasal endoscopy and MRI. A total of 8,335 subjects were enrolled by September 15, 2018. Two rounds of screening were completed: 784 (9.4%) subjects had plasma EBV Retesting four weeks later revealed that 230 subjects (2.7%) had positive DNA. However, all patients had detectable plasma EBV DNA. The test results from both rounds are summarized below. [Table 1]

[0141] As shown in Table 1, plasma EBV DNA was detected in the second round of NPC screening. The probability of receiving EBV correlated with plasma EBV DNA status at the first round of screening. The first round of screening showed negative, transiently positive, and persistently positive results. Subjects with plasma EBV DNA were identified in the first analysis of the second round of screening. The probability of having extractable plasma EBV DNA was 8%, 21%, and 57%. The chance of persistent plasma EBV DNA positivity at 4 weeks was 1.0% across the three groups. It gradually increased from 2% to 25%.

[0142] NPC patients identified by the screening described in this specification showed a much earlier stage distribution than patients in a past cohort who did not undergo NPC screening. That is. The percentages of early-stage diseases (stages I and II) were 70% and 20% respectively. respectively. Due to this change in the disease stage distribution, the progression-free survival period of patients with a hazard ratio of 0.1 was greatly improved. Summarized in Table 2 is the disease stage distribution of NPC cases in both the first and second rounds of screening. After screening 8,335 subjects in the second round, 13 new NPC cases were identified. The percentages of patients with early-stage diseases were 71% and 69% respectively in the first and second rounds of screening. There was no significant difference in the percentage of patients with early-stage diseases (P = 0.93, chi-square test).

Table 2

[0143] As summarized in Table 3, subjects with temporarily and persistently detectable plasma EBV DNA in the first round of screening had a higher risk of NPC detection in the second round of screening conducted 4 years after the first round compared to those in whom plasma EBV DNA was not detected in the first round. The relative risk values for these two groups were 7.2 and 19.7 respectively.

Table 3

[0144] These results indicate that plasma EBV DNA analysis is a screening for the current state of NPC prevalence​​​​​​​​ suggesting that it is also useful for predicting the risk of NPC that can be clinically observed in the future, not only for the ring. One practical application of this finding is that the interval for repeating the screening can be adjusted based on the plasma EBV DNA status of the subjects screened in previous examples. That is, for example, subjects who have detectable plasma EBV DNA at baseline but no NPC identified can be re-screened at shorter intervals compared to subjects with undetectable plasma EBV DNA. Also, as an example, the intervals for repeating the screening can be 4 years, 2 years, and 1 year for subjects with undetectable, temporarily detectable, and continuously detectable plasma EBV DNA, respectively.

[0145] [Example 2. NPC Screening Based on Detectability of Plasma EBV DNA] This example describes an NPC screening regimen designed for subjects based on the detectability of EBV DNA in the plasma of the subjects. Figure 2 shows a schematic diagram of the regimen described herein.

[0146] According to the regimen, subjects with undetectable plasma EBV DNA in the initial example of the screening have a relatively low risk of NPC of subjects with undetectable EBV DNA in the next four years, so they are re-screened after 4 years. If the plasma EBV DNA is negative in the next screening, the interval for the next screening is 4 years. However, if EBV DNA is detected in one screening but no NPC is detected, the next screening is adjusted to one year later. If the plasma EBV DNA remains negative for 4 years, the screening The screening interval is reset to 4 years. The actual time intervals used in a particular screening program are also adjusted according to healthcare economic considerations (such as the cost of screening), subject preferences (e.g., being more frequent than the screening interval may cause more disruption to a particular subject's lifestyle), and other clinical parameters (e.g., an individual's genotype, family history of NPC, diet history, ethnic origin (e.g., Cantonese)).

[0147] [Example 3. Mutation Pattern Analysis of Cell-Free EBV DNA Molecules] In this example, targeted sequencing with capture enrichment was used to analyze cell-free viral DNA molecules in the circulation of NPC subjects, non-NPC subjects with detectable plasma EBV DNA, and pre-NPC subjects (details in the next section). The capture probes were designed to cover the entire EBV genome. In this analysis, probes targeting approximately 3,000 human single nucleotide polymorphism (SNP) sites and human leukocyte antigen (HLA) SNPs were also included.

[0148] In this example, plasma EBV DNA of 13 NPC patients and 16 non-NPC subjects with detectable plasma EBV DNA was analyzed. The 13 NPC patients were symptomatic and recruited from either the Clinical Oncology Department or the Otolaryngology Department of the Prince of Wales Hospital. The 16 non-NPC subjects were from the NPC screening cohort of over 20,000 subjects described in Example 1.

[0149] In this analysis, targeted sequencing with capture enrichment by specially designed capture probes Cincin was used. For each plasma sample analyzed, DNA was extracted from 4 mL of plasma using the QIAamp Circula ting Nucleic Acid Kit. In all cases, all the extracted DNA was used for the preparation of sequencing libraries using the TruSeq Nano DNA l ibrary preparation kit (Illumina). Bar coding was performed using a dual-index system (xGen dual-index UMI adapter, Integrated DNA Techn ologies) that incorporated a unique molecular identifier (UMI) sequence. 8-cycle PCR amplification was performed on samples ligated with adapters using the TruSeq Nano K it (Illumina). Next, the amplification products were captured by the myBait custom capture panel system (Arbor Bioscien ces) using custom-designed probes covering the viral and human genomic regions described above. After target capture, the captured products were concentrated by 14-cycle PC R to generate a DNA library. The DNA library was sequenced on the NextSeq platform (Illumina). For each sequencing run, 10 samples with unique sample barcodes were sequenced using the paired-end mode. Each D NA fragment was 71 nucleotides sequenced from each of the two ends. After sequencing, the sequence reads were mapped to an artificially combined reference sequence consisting of the entire human genome (hg19), the entire EBV genome (GenBank: AJ507799.2), the entire HBV genome, and the entire HPV genome. The alignment was SOA R R to generate a DNA library. The DNA library was sequenced on the NextSeq platform (Illumina). For each sequencing run, 10 samples with unique sample barcodes were sequenced using the paired-end mode. Each D NA fragment was 71 nucleotides sequenced from each of the two ends. After sequencing, the sequence reads were mapped to an artificially combined reference sequence consisting of the entire human genome (hg19), the entire EBV genome (GenBank: AJ507799.2), the entire HBV genome, and the entire HPV genome. The alignment was SOA R R to generate a DNA library. The DNA library was sequenced on the NextSeq platform (Illumina). For each sequencing run, 10 samples with unique sample barcodes were sequenced using the paired-end mode. Each D NA fragment was 71 nucleotides sequenced from each of the two ends. After sequencing, the sequence reads were mapped to an artificially combined reference sequence consisting of the entire human genome (hg19), the entire EBV genome (GenBank: AJ507799.2), the entire HBV genome, and the entire HPV genome. The alignment was SOA R R to generate a DNA library. The DNA library was sequenced on the NextSeq platform (Illumina). For each sequencing run, 10 samples with unique sample barcodes were sequenced using the paired-end mode. Each D Performed using P2 (Bioinformatics 2009; 25:1966-7), allowing a maximum of two mismatches each time a correct reading was taken in the forward direction with an insert size of 600 bp or less. Sequence reads mapped to unique positions in the assembled genomic sequences were used for downstream analysis. All overlapping fragments with exactly the same unique molecular identifier were filtered out. Based on the alignment results, nucleotide differences including, but not limited to, single nucleotide polymorphisms (SNVs) were identified between the sequenced reads and the EBV reference genome (GenBank: AJ507799.2). Among 44 samples from 13 NPC subjects, 16 non-NPC subjects with detectable plasma EBV DNA, and 4 pre-NPC subjects, a median of 1116 SNVs (interquartile range (IQR): 902-1216) were identified. In these plasma samples, two different alleles were observed at several nucleotide positions in the EBV genome. This observation could be due to sequence errors or the presence of tumor heterogeneity. A median of only 26 positions (IQR: 20-35) had two or more alleles in plasma EBV DNA. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished.

[0150] Based on the alignment results, nucleotide differences including, but not limited to, single nucleotide polymorphisms (SNVs) were identified between the sequenced reads and the EBV reference genome (GenBank: AJ507799.2). Among 44 samples from 13 NPC subjects, 16 non-NPC subjects with detectable plasma EBV DNA, and 4 pre-NPC subjects, a median of 1116 SNVs (interquartile range (IQR): 902-1216) were identified. In these plasma samples, two different alleles were observed at several nucleotide positions in the EBV genome. This observation could be due to sequence errors or the presence of tumor heterogeneity. A median of only 26 positions (IQR: 20-35) had two or more alleles in plasma EBV DNA. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished.

[0151] In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished. In the phylogenetic analysis shown in Figure 3, NPC subjects clustered together and separated from non-NPC subjects. These results suggested that there were different EBV variant profiles between NPC and non-NPC subjects. Therefore, by using EBV variant profile analysis of plasma EBV DNA, subjects with NPC and non-NPC in the context of screening could be distinguished. ​can be identified. Three non-NPC subjects (AC106, AP080, and FF 159) had two consecutively collected and analyzed samples collected at 4-week intervals . Clustering two samples from the same individual showed that they shared highly similar variants .

[0152] Phylogenetic analysis was also performed on the same group of 13 NPC patients and 16 non-NPC subjects with detectable plasma EBV DNA, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). As shown in Figure 4, NPC subjects also clustered together and were separated from non-NPC subjects . Four subjects who were persistently positive for plasma EBV DNA in the first round of screening (as described in Example 1) but had no detectable NPC by endoscopy and MRI were subsequently diagnosed with NPC. All of them (BB0

[0153] 96, DN054, FK015, and HB121) were diagnosed with NPC 3 years after the first round of screening . All of them had one additional plasma sample collected 1 year after the first round of screening during follow-up at the otolaryngology clinic . For each of these four subjects, two samples collected at the first round of screening and 1 year later were analyzed for EBV variants . As shown in Figure 5, samples from pre-NPC subjects clustered with NPC samples . ​​​​The EBV variants associated with this have been shown to exist prior to the actual development of cancer. This indicates that individuals with NPC-related EBV variants have a high risk of developing NPC in the future. Phylogenetic analysis also excluded the 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049), and was also performed on the same groups of NPC patients, non-NPC subjects, and pre-NPC patients based on the EBV variants excluding those 29 variants. As shown in Figure 6, samples from pre-NPC subjects were also clustered together with NPC samples, further suggesting that the risk of future NPC can be predicted by the analysis of EBV variants.

[0154] [Example 4. Block-based Mutation Pattern Analysis] This example describes the working principle of an exemplary block-based variant pattern analysis approach and its application to the analysis of EBV variant patterns within samples described in Example 3.

[0155] Figure 7 illustrates the principle of block-based mutation pattern analysis. Using block-based analysis, the similarity of EBV DNA mutation patterns derived from plasma EBV DNA sequencing of various samples to the reference genome is evaluated, and here, NPC sequencing data available in public databases (Kwok et al. J Virol 2014; 88:10662-72, Li et al. Nat Comm 2017; 8:1 4121) is used as a reference. In block-based analysis, the EBV genome is divided into bins of 500 bp in size (a total of 344 bins), and each bin The similarity between the mutation pattern and 24 NPC samples of the reference set was compared. As an example , if there are 8 mutation sites in one specific bin, the alleles of these sites in this bin of the test sample are analyzed and compared with the alleles of the same sites in 24 reference samples. The similarity index is derived based on the proportion of having exactly the same alleles as the reference samples. For example, if the test sample has exactly the same alleles as 7 out of 8 mutation sites for one reference sample, the similarity index of that bin for that reference sample is 7 / 8. Also, when comparing with 24 reference samples, there are 24 similarity indices for that bin of the test sample. Based on the 24 similarity indices of that bin, a bin score representing the overall similarity of the mutation pattern for the reference sample is calculated. For example, when setting the cut-off of the similarity index to 0.9, the bin score counts the proportion of bins having an index higher than the cut-off. Therefore, if only 2 out of 24 similarity indices exceed 0.9, the bin score is 2 / 24. The higher the bin score, the more similar the mutation pattern of the test sample is to the reference sample set. Figure 8 shows a block-based analysis of the EBV DNA mutation patterns of 13 NPCs, 16 non-NPCs, and 4

[0156] pre-NPC samples. For each of the 4 pre-NPC subjects, samples from 2 time points were analyzed, resulting in a total of 8 subjects. The bin scores of 344 bins of the EBV genome were derived from these samples. Based on the bin scores of these samples, an unsupervised clustering analysis was performed. For the NPC samples ​Both the NPC samples (black) and non-NPC samples (marked with dots) were clustered together. Both were clustered together. The EBV mutation profiles of pre-NPC subjects were clustered together with those of NPC subjects. In particular, for the mutation profiles of these four pre-NPC subjects, they were obtained through the analysis of baseline samples collected several years before the onset of NPC. Both were clustered together. The EBV mutation profiles of pre-NPC subjects were clustered together with those of NPC subjects. In particular, for the mutation profiles of these four pre-NPC subjects, they were obtained through the analysis of baseline samples collected several years before the onset of NPC. Both were clustered together. The EBV mutation profiles of pre-NPC subjects were clustered together with those of NPC subjects. In particular, for the mutation profiles of these four pre-NPC subjects, they were obtained through the analysis of baseline samples collected several years before the onset of NPC. Both were clustered together. The EBV mutation profiles of pre-NPC subjects were clustered together with those of NPC subjects. In particular, for the mutation profiles of these four pre-NPC subjects, they were obtained through the analysis of baseline samples collected several years before the onset of NPC.

[0157] Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients. Figure 9 shows a block-based analysis of EBV DNA variants of the same group of 13 NPC, 16 non-NPC, and 4 pre-NPC subjects, based on EBV variants excluding 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc.32049). Similarly, clustering of NPC samples (black) was observed. Also, the EBV variant profiles of pre-NPC subjects were clustered together with those of NPC subjects. The clustering of pre-NPC samples and NPC samples indicates that mutation analysis can predict the future onset of NPC. In summary, the data of Example 3 and Example 4 show that subjects who did not have NPC at recruitment but later developed cancer had EBV mutation patterns in baseline blood samples similar to those from other NPC patients.

[0158] [Example 5. Risk Prediction of NPC Using a Mathematical Model] This example describes the construction of a classification model for predicting the future risk of NPC onset in subjects with detectable plasma EBV DNA using mutation pattern analysis, and the test results using the classification model. This example describes the construction of a classification model for predicting the future risk of NPC onset in subjects with detectable plasma EBV DNA using mutation pattern analysis, and the test results using the classification model. This example describes the construction of a classification model for predicting the future risk of NPC onset in subjects with detectable plasma EBV DNA using mutation pattern analysis, and the test results using the classification model.

[0159] Using the support vector machine (SVM) algorithm, as described in Example 4, a classifier was constructed using the training dataset of 18 subjects without NPC and 8 NPC patients. The test dataset consisted of 5 NPC patients, 5 subjects without NPC, and 8 samples collected from 4 subjects who had no detectable NPC by endoscopy and MRI during sample collection as described in Example 4 but were later diagnosed with NPC (labeled pre-NPC). The method of SVM analysis is as follows: When a training dataset containing n samples is given: (M1, Y1), …, (Mn, Yn) where Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC.

[0160] By finding a set of coefficients (W with a p-dimensional vector) that satisfy the following, a "hyperplane" can be identified that separates the non-NPC group and the NPC group as accurately as possible within the training dataset: Criterion 1: W·M Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC. Here, Yi indicates the NPC status of sample i. Yi is 1 for samples from NPC patients or -1 for samples from subjects without NPC; Mi is a p-dimensional vector containing the viral mutation pattern of sample i. For example, Mi can be a series of mutation sites such as 29 mutation sites related to NPC. Alternatively, Mi can be a series of block-based variant similarity scores (e.g., non-overlapping windows of 500 bp) for a reference EBV variant present in subjects known to have NPC.

[0161] By finding a set of coefficients (W with a p-dimensional vector) that satisfy the following, a "hyperplane" can be identified that separates the non-NPC group and the NPC group as accurately as possible within the training dataset: By finding a set of coefficients (W with a p-dimensional vector) that satisfy the following, a "hyperplane" can be identified that separates the non-NPC group and the NPC group as accurately as possible within the training dataset: By finding a set of coefficients (W with a p-dimensional vector) that satisfy the following, a "hyperplane" can be identified that separates the non-NPC group and the NPC group as accurately as possible within the training dataset: Criterion 1: W·Mi -b ≥ 1 (for subjects in the NPC group) and Criterion 2: W·M i -b ≤ 1 (for subjects not in the NPC group) Here, W is a p - dimensional vector of coefficients that determines the hyperplane; M is a matrix (p x n order) with p mutants ( or block - based similarity scores) and n samples; b is the intercept.

[0162] The two criteria (i.e., Criterion 1 and 2) can also be described as follows: Yi (W * Mi - b) ≥ 1 (Criterion 3) Here, Yi is either - 1 (non - NPC) or 1 (NPC).

[0163] The margin distance (D) between Criterion 1 and 2 is:

Number

[0164] D is maximized by minimizing JPEG2025084804000013.jpg923 according to Criterion 3.

[0165] Based on this principle, the parameters (W and b) of the classifier were determined. Then, using the trained parameters (W and b), the NPC risk score for each test sample was calculated.

[0166] Figure 10A shows the NPC risk scores calculated using a classifier trained based on the analysis of all EBV mutants using block - based variant analysis. In this analysis, as described in Example 4, the EBV genome was divided into 500bp bins to calculate the bin scores ​It was divided into 44 blocks. The bin score was regarded as a feature of machine learning. NPC samples The NPC risk scores of were significantly higher than those of samples collected from non-NPC subjects (average NPC risk score: 0.15 vs. 0.53, p-value < 0.01, Student's t-test). Similarly, the NPC risk scores were significantly higher in samples collected from pre-NPC subjects compared to subjects without NPC (average risk score: 0.5 8 vs. 0.15, p-value < 0.01, Student's t-test). Using a cutoff of 0.32, samples from NPC patients and pre-NPC subjects could be discriminated from samples without NPC with 100% sensitivity and 1 00% specificity. Figure 10B shows the NPC risk scores calculated using a classifier trained based on the analysis of 29 variants reported in the study by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc .32049). The NPC risk scores of NPC samples

[0167] were significantly higher than those of samples collected from non-NPC subjects (average NPC risk score: 0.89 vs. 0.18, p-value < 0.01, Student's t-test). Similarly, the NPC risk scores were significantly higher in samples collected from pre-NPC subjects compared to subjects without NPC (average risk score: 0.57 vs. 0.18, p-value <0.02, Student's t-test). Using a cutoff of 0.6, samples from NPC patients and pre-NPC subjects could be discriminated from samples without NPC with 74% sensitivity and 100% specificity.

[0168] ​​​​​​​​​ Figure 10C shows the NPC risk scores calculated using a classifier trained based on the analysis of all EBV variants using a bulk-based variant analysis, excluding 29 variants previously reported to be associated with NPC by Hui et al (Hui et al. Int J Cancer 2019, doi.org / 10.1002 / ijc .32049). The NPC risk scores of NPC samples were significantly higher than those of samples collected from non-NPC subjects (mean NPC risk score: 0.58 vs. 0.15, p-value < 0.01, Student's t-test). Similarly, the NPC risk scores were significantly higher in samples collected from pre-NPC subjects compared to subjects without NPC (mean risk score: 0.53 vs. 0.15, p-value < 0.01, Student's t-test). Using a cut-off of 0.31, samples from NPC patients and patients who subsequently developed NPC could be discriminated from samples without NPC with 100% sensitivity and 100% specificity. These results indicate that excluding the previously reported 29 EBV variants from the analysis does not adversely affect the accuracy of this analysis. Example 6. Analysis of the Methylation Status of Plasma EBV DNA by Bisulfite Sequencing This example demonstrates the use of bisulfite sequencing to discriminate NPC patients and non-NPC subjects with detectable plasma EBV DNA based on the methylation status of plasma EBV DNA. The methylation levels of EBV DNA in the plasma of NPC patients and subjects without NPC

[0169]

[0170] ​​​​​​​​​​​ was determined using bisulfite sequencing. Bisulfite conversion can change unmethylated cytosine to uracil. Methylated cytosine cannot be changed by bisulfite and can remain as cytosine. During sequence determination, uracil can be determined as thymine. After sequence determination, by checking whether cytosine has changed to thymine, the methylation status of cytosine in any CpG dinucleotide context can be determined. The methylation levels of plasma EBV DNA were determined in 10 NPC patients and 40 subjects (non-NPC subjects) who did not have cancer but had detectable EBV DNA in their plasma . For the 40 non-NPC subjects, another blood sample was collected from each of them 4 weeks later . Twenty of them became negative for plasma EBV DNA and they were labeled as having transiently positive plasma EBV DNA

[0171] . Twenty of them remained positive for plasma EBV DNA and they were labeled as having persistently positive plasma EBV DNA . As shown in Figure 11, the EBV DNA methylation levels were significantly higher in NPC patients compared to non-cancer subjects with transiently positive plasma EBV DNA (p value < 0.01, Student's t-test) and non-cancer subjects with persistently positive plasma EBV DNA (p value < 0.01, Student's t-test ) . These results suggest that the analysis of plasma EBV DNA methylation may be useful for discriminating between NPC patients and subjects who do not have NPC but have detectable plasma EBV DNA .

[0172] . ​

[0173] [Example 7. Analysis of the Methylation Status of Plasma EBV DNA Using Methylation-Sensitive Restriction Enzymes This example demonstrates the use of methylation-sensitive restriction enzyme analysis of plasma EBV DNA for the identification of NPC patients and subjects without NPC but with detectable plasma EBV DNA, and describes an in-silico simulation experiment.

[0174] Bisulfite sequencing of plasma DNA was performed on samples from non-NPC subjects and NPC patients. 347,516 and 6, 271,012 EBV DNA fragments were obtained in the plasma DNA of two subjects, respectively. Their plasma EBV DNA methylation levels were 48.9% and 86.3%, respectively. It was determined that approximately half of the plasma EBV D NA molecules contain at least one "CCGG" motif.

[0175] To simulate the restriction enzyme digestion of plasma EBV DNA, in-silico digestion of plasma EBV DNA molecules was performed according to the methylation status in the "CCGG" sequence context inferred from the results of bisulfite sequencing. Thus, as shown in Figure 14, simulated size profiles of plasma EBV DNA were obtained with and without in-silico digestion by the methylation-sensitive restriction enzyme HpaII. In the absence of enzyme digestion, the size distribution of plasma EBV DNA in non-NPC subjects was to the left of that in NPC subjects, indicating that the size distribution was shorter in non-NPC subjects. This difference in fragment size indicates that the non-NPC subjects with enzyme digestion had a different size distribution compared to those without enzyme digestion. ​​​​​​​​​​​In the subjects, an increase in the abundance of short DNA less than 50 bp was significant, and this was also observed in the size distribution profile with enzyme digestion. For NPC patients, the proportion of DNA molecules less than 50 bp was 5.87 % and 0.84% in the samples with and without enzyme digestion, respectively. However, for non-NPC subjects, the proportion of DNA molecules less than 50 bp was 22.24% and 4.99% in the samples with and without enzyme digestion, respectively. The increase in the proportion of DNA less than 50 bp in enzyme digestion was 17.2% and 5.0% in NPC patients and non-NPC subjects, respectively. Figure 15 shows the cumulative size profiles of plasma EBV DNA with and without methylation-sensitive restriction enzyme digestion for NPC patients and non-NPC subjects. The difference in the degree of enzyme digestion can be more easily understood using the cumulative frequency curve against size. The gap between the two curves with and without enzyme digestion reflects the degree of digestion. The larger the gap, the greater the degree of enzyme digestion performed on plasma EBV DNA, indicating a low methylation level of plasma EBV DNA. As shown in the figure, the non-NPC subjects had a larger gap compared to NPC patients. For NPC patients and non-NPC subjects, the maximum distance between the curves without and with enzyme digestion was 8.1 and 18.3, respectively; for NPC patients and non-NPC subjects, the area between the two curves was 2395 and 942.9, respectively. respectively. The difference in the degree of enzyme digestion is more easily understood using the cumulative frequency curve against size. The gap between the two curves with and without enzyme digestion reflects the degree of digestion. The larger the gap, the greater the degree of enzyme digestion performed on plasma EBV DNA, indicating a low methylation level of plasma EBV DNA. As shown in the figure, the non-NPC subjects had a larger gap compared to NPC patients. For NPC patients and non-NPC subjects, the maximum distance between the curves without and with enzyme digestion was 8.1 and 18.3, respectively; for NPC patients and non-NPC subjects, the area between the two curves was 2395 and 942.9, respectively. Compared to NPC patients, non-NPC subjects had a larger gap. For NPC patients and non-NPC subjects, the maximum distance between the curves without and with enzyme digestion was 8.1 and 18.3, respectively; for NPC patients and non-NPC subjects, the area between the two curves was 2395 and 942.9, respectively. respectively.

[0176] [Example 8. SNV Profile Analysis of Cell-Free EBV DNA Molecules] A total including plasma DNA sequencing data of 63 NPC and 88 non-NPC subjects In the training dataset, the differences in EBV SNV profiles between two groups were analyzed. Identifying SNVs across the EBV genome was identified. The NPC risk score is to be derived from the genotype patterns of these SNV sites, and then , it was analyzed in a test set of 31 NPC samples and 40 non-NPC samples . In this example, a total of 661 important SNVs across the entire EBV genome were identified from the training set (Figure 16D). In the test set, NPC plasma samples were shown to have a high NPC risk score; an NPC-related EBV SNV profile may exist. Among non-N PC samples, the NPC risk scores were widely distributed. Non-NPC subjects can have diverse EBV SNV profiles.

[0177] Materials and methods.

[0178] Study participants and design.

[0179] This study included a subset of the sequencing dataset of NPC and non-NPC plasma samples previously reported by Lam et al. Proc Natl Acad Sci U S A. 2018; 115:E5115-E5124 (as the training set), and the analysis of newly sequenced plasma DNA samples from both NPC and non-NPC subjects (as the test set). The training dataset is from past prospective NPC screening studies described in Lam et al. Proc Natl Acad Sci U S A. 2018; 115:E5 115-E5124.

[0180] 115-E5124. Plasma samples were included from both NPC patients and non - NPC subjects detected by screening. These non - NPC subjects harbored plasma EBV DNA at detectable levels by real - time PCR - based assays. This dataset also included samples from symptomatic NPC patients from an independent cohort. To construct a training model for NPC risk score prediction, EBV genotype information from EBV isolates of all samples was studied. In this study, plasma samples from another 31 symptomatic NPC patients and 40 non - NPC subjects were targeted for target capture sequencing as a test set. These 31 symptomatic NPC patients were recruited from the Department of Clinical Oncology, Prince of Wales Hospital, Hong Kong. Non - NPC subjects were also from the aforementioned NPC screening cohort (including over 20,000 subjects) and were randomly selected from there. The variation of EBV genotypes from these NPC and non - NPC samples was analyzed and their NPC risk scores were derived based on the training model. All NPC samples and non - NPC samples in the training set and test set were non - overlapping.

[0181] Target capture sequencing

[0182] Target capture sequencing of plasma samples was performed by enriching EBV DNA molecules from plasma DNA libraries via a capture probe system (myBaits custom capture panel, Arbor Biosciences). EBV capture probes were designed to cover the entire viral genome. 3,000 human single - nucleotide polymorphism (SNP) sites were targeted ​ Probes to be used are also included for reference. A probe mixture containing the EBV probe and the autosomal DNA probe in a molar ratio of 100:1 was used in each capture reaction. DNA libraries from 10 plasma samples were multiplexed in one capture reaction without using the same amount of DNA library from each sample. Sequencing statistics for all cases, including previously reported cases used as the current training set, are described in Tables 4A and 4B.

[0183]

Table 4A-1

Table 4A-2

Table 4A-3

Table 4A-4

Table 4A-5

[0184]

Table 4B-1

Table 4B-2

[0185] EBV variant calling

[0186] The sequenced reads were aligned to the human (hg19) and EBV reference genomes using the BWA aligner described in Li H et al. Bioinformatics. 2010; 26:589-95. (AJ507799.2), which is incorporated herein by reference in its entirety. Alternative alleles that differ from the reference viral genome on the EBV genome site are incorporated into the text. If EBV single nucleotide polymorphisms (SNVs) were detected in the offspring, Li H et al. Bioinformatics. 2009; 25:2078-9, which was identified using Samtools, The present invention is incorporated herein by reference in its entirety. V sites (with minor allele frequency cutoff set at 5%), followed by were excluded for the NPC risk score analysis.

[0187] NPC Risk Score

[0188] In this example, the NPC risk score is calculated based on the SNV sites across the viral genome. Weighted sum of EBV genotypes in a fixed set (explanation variables of binary logistic regression model) The set of NPC-associated SNVs was analyzed using the NPC and First, by analyzing differences in EBV SNV profiles from non-NPC samples Using Fisher's exact test, we identified EBV genome sequence for NPC cases. The association of each variant across the cohort was then analyzed. The false discovery rate (FDR) was then controlled at 5% to determine the significance of each variant. We obtained a fixed set of significant SNVs.

[0189] The NPC risk score of the test sample was calculated based on the significant scores identified from the training set. This can be determined by the EBV genotype for this particular set of NV sites. As mentioned above, the concentration of plasma EBV DNA molecules is low, so the sequenced EBV DNA With the means by reads, coverage of the entire EBV genome may be incomplete. Thus the score was formulated to be determined by the genotype pattern across those SNVs covered by plasma EBV DNA reads (e.g., using available genotype information) (FIGS. 16A, 16B, and 16C). To derive the NPC risk score, a subset of important SNV sites was first identified and covered by plasma EBV DNA reads of the test sample. Then, the weighting ( effect size) of the genotype at each site was determined within the subset of important SNV sites. This was done by analyzing the genotype patterns at each site between NPC samples and non-NPC samples in the training dataset (FIG. 16B). Based on this, a logistic regression model was constructed and provided information on the effect size of the risk genotype at each SNV site of NPC. The logistic model was described as follows:

Equation

Equation

[0190] Results

[0191] Construction of the NPC risk score training model

[0192] As described above, the previously reported plasma EBV DNA sequencing data of NPC and non-NPC samples was used for the development of the NPC risk score training model. To concentrate EBV DNA in the plasma sample, target capture sequencing was performed Performed. The viral SNV profiles of EBV isolates from NPC and non-NPC samples files were studied here. From this dataset, EBV DNA reads that were sequenced had at least 30% coverage of the entire EBV genome by NPC and non-NP C cases were selected. This cutoff was chosen because more than 95% of the NPC samples in the training dataset had viral genome coverage greater than the cutoff (Table 4A and 4B). The demographics of these selected NPC and non-NPC subjects, including age and gender, and the cancer stage information of NPC patients (8th AJC C edition) are shown in Table 5. The sequencing statistics of these selected NPC and non-NPC samples are described in (Table 4A and 4B).

Table 5

[0193] The EBV SNV profiles of these 63 NPC samples and 88 non-NPC samples were analyzed. The median sequence depth across the entire EBV genome for all samples was 2-fold (interquartile range (IQR), 1.0-fold to 9.2-fold). The average number of EBV SNVs identified from NPC samples was 800 (IQR, 662 - 958), and the average number of SNVs among non-NPC samples was 539 (range, 363 - 656). In total, 5678 different SNVs were identified in all samples. The distribution of these SNVs across the EBV genome is shown in Figure 16.

[0194] The association of each viral SNV with NPC samples in the training set was also determined by Fisher - was studied by direct probability testing. By controlling the false discovery rate (FDR) to 0.05, , a total of 661 significant SNVs associated with NPCs having adjusted p-values were identified. These genomic positions of the 661 SNVs are shown in Table 6. Subsequently, the NPC risk scores of the test set of plasma samples of NPC and non-NPC subjects were derived based on the genetic pattern of these 661 SNV sites.

Table 6-1

Table 6-2

Table 6-3

[0195] Evaluation of the NPC Risk Score Training Model

[0196] The training model was evaluated to analyze the NPC risk scores of samples within the training set using the leave-one-out approach. In the leave-one-out approach, the principles for constructing the training model and deriving the NPC risk scores were the same as those described by the method. All but one sample of the training set were used to construct the training model, and the excluded sample could be analyzed for NPC risk scores. In the leave-one-out approach, the median NPC risk score of the NPC group was 0.99 (IQR, 0.98 - 1.0), and the median of the non-NPC group was 0.01 (IQR, 0.00 - 0.89) (Figure 17A). Receiver operating characteristic (RO C) curve analysis was used to analyze NPC samples and non-NPC samples by NPC risk scores The discrimination with [object] was evaluated. The area under the curve value was 0.91 (Figure 17B).

[0197] NPC risk score analysis in the test set

[0198] The target capture sequence was performed on plasma samples from another 31 NPC patients and 45 non-NP C subjects. Among them, 31 NPC samples and 40 non-NPC samples all had at least 30% or more coverage of the EBV genome by the sequenced EBV DNA reads. The clinical characteristics of these NPC and non-NPC subjects are summarized in Table 7. The sequencing statistics of this series of test samples are also described in Tables 4A and 4B as well.

Table 7

[0199] Based on the developed training model, the NPC risk scores of the test sets of 31 NPC samples and 40 non-NP C samples were analyzed. The NPC risk score of a sample can be determined by its mutation pattern across 661 important SNV positions identified from the training set. Due to the possible incomplete coverage of the EBV genome, only SNV sites covered by the sequenced EBV DNA reads and having corresponding allele information can be included in the NPC risk score analysis (Figures 16A, 16B and 16C).

[0200]

[0200] The median NPC risk score of the NPC group was 0.999 (IQR, 0.996 - 0 .999), and that of the non-NPC group was 0.557 (IQR, 0.000 - 0. It was 996 (Figure 18A). Similarly, among these 31 NPC samples, high NP C risk scores were observed. The NPC samples in the test set can share the same EBV SNV profile as the NPC samples in the training set . The identification of NPC and non-NPC samples by the NPC risk score was also evaluated by ROC curve analysis . The area under the curve value was 0.83 (Figure 18B).

[0201] Analysis of genotype patterns across high-risk variant sites in the test set

[0202] In the EBER (small RNA encoded by EBV) region, there are EBV variants related to high-risk NPC . In the EBER region, 23 important SNVs have been reported by Hui et al . A similar approach for NPC risk prediction was adopted in a test set of 31 NPC samples and 40 non-NPC samples, but the analysis was based only on the genotype patterns of 2 3 of the SNVs reported in the EBER region .

[0203] In the test set, 31 (44%) of the 71 NPC and non-NPC samples had EBV DNA reads covering all 2 3 SNV sites. As shown in Table 8, for each of these 23 SNV sites, only a portion of the samples had available genotype information containing reads covering the SNV site (i.e , not all 23 SNV sites were covered by the plasma EBV DNA reads of the samples ). The percentage of high-risk genotypes at each of the 23 SNV sites in NPC samples ranges from 86% to 97%. The high-risk genes in non-NPC samples type percentage...​​​ The percentage of the [[TYPE]] is in the range of 35% to 52%. Analyzed NPC and non-NPC The number of samples extends to samples containing available genotype information (e.g., including EBV DNA reads covering the SNV site ). In the test set (31 NPC samples and 40 non-NPC samples), only some samples had reads covering the SNV site and genotype information available at the corresponding sites. The discrimination between NPC samples and non-NPC samples was also evaluated only by analyzing the genotype patterns of 23 SNVs in the EBER region by ROC curve analysis. The area under the curve was 0.72 (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. (FIGS. 19A and 19B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. Analysis of genotype patterns across the entire EBV genome can achieve better discrimination between NPC samples and non-NPC samples than analysis across fixed viral genome regions. [Table 8]

[0204] Similarly, three high-risk SNVs in the BALF2 (BamHI A left frame-2) gene have also been reported (Xu et al. Nat Genet. 2019; 51:1131-6). In the test set, 55 (78%) of 71 samples had EBV DNA reads covering all three SNVs. For each of these three SNV sites, only some samples in the test set had reads covering the SNV site with available genotype information (Table 9). The high-risk genotypes at each of the three SNV sites in NPC samples (Table 9). The high-risk genotypes at each of the three SNV sites in NPC samples (Table 9). The high-risk genotypes at each of the three SNV sites in NPC samples (Table 9). The high-risk genotypes at each of the three SNV sites in NPC samples The percentage of the subtype ranges from 86% to 93%. The percentage of high-risk genotypes in non-NPC samples ranges from 47% to 65%. There were 4 cases without EBV DNA reads covering any of the 3 reported SNVs (1 NPC sample and 3 non-NPC samples), and these cases could not be analyzed. The same approach for NPC risk prediction was adopted for the remaining 30 NPCs and 37 non-NPC samples from the test set, and only the genotype patterns of the 3 SNVs reported in the BALF2 region were analyzed. The discrimination between NPC samples and non-NPC samples was also evaluated by ROC curve analysis. The area under the curve was 0.77 (Figures 20A and 20B). This value was lower than the value (0.83) obtained from the analysis of genotype patterns across the entire EBV genome. The analysis of genotype patterns across the entire EBV genome achieved better discrimination between NPC samples and non-NPC samples than the analysis across fixed viral genome regions.

Table 9

[0205] In the NPC risk score analysis described in this example, NPC risk prediction based on the genotype patterns across the floating number of SNVs randomly selected within a set of 661 important SNVs on the EBV genome is possible (Table 6). The floating number of SNV sites used for NPC risk score analysis is determined by whether the SNV sites are covered by the EBV DNA reads for which the sequences were determined and have corresponding allele information. is performed. Downsampling of the set of 661 important SNVs is carried out, and the performance of the sample NPC prediction is analyzed in the test set using the same approach as the floating number of SNVs in the downsampled set of SNVs. In the downsampling analysis, a certain number (such as 23, 25, 100, 200, or 500, etc.) of SNVs are randomly selected from the 661 important SNVs. Then, for the test sample, the SNV sites within the downsampled set of SNVs covered by the EBV DNA sequence reads are identified. Then, in the training set of the covered and downsampled SNV sites, the NPC risk score training model is obtained by training the model using the genotype patterns of NPC samples and non - NPC samples. Through this training, the weight of the genotype at each site for the training model is determined. And by applying its own genotype pattern across these covered and downsampled SNV sites to the NPC risk score training model weighted by the downsampled SNV sites, the NPC risk score of the test sample is derived. The prediction performance of the NPC risk score training model with various numbers of SNV sites is summarized in Table 10. For a given number of SNV sites, SNVs are randomly selected for downsampling 10 times, and the area under the curve value in Table 10 is the average result of the 10 - time random downsampling. The set of SNVs across the entire EBV genome is downsampled to 23, which is the same as the number of SNVs reported in the EBER region. NPC sample is analyzed in the test set using the same approach as the floating number of SNVs in the downsampled set of SNVs. In the downsampling analysis, a certain number (such as 23, 25, 100, 200, or 500, etc.) of SNVs are randomly selected from the 661 important SNVs. Then, for the test sample, the SNV sites within the downsampled set of SNVs covered by the EBV DNA sequence reads are identified. Then, in the training set of the covered and downsampled SNV sites, the NPC risk score training model is obtained by training the model using the genotype patterns of NPC samples and non - NPC samples. Through this training, the weight of the genotype at each site for the training model is determined. And by applying its own genotype pattern across these covered and downsampled SNV sites to the NPC risk score training model weighted by the downsampled SNV sites, the NPC risk score of the test sample is derived. The prediction performance of the NPC risk score training model with various numbers of SNV sites is summarized in Table 10. For a given number of SNV sites, SNVs are randomly selected for downsampling 10 times, and the area under the curve value in Table 10 is the average result of the 10 - time random downsampling. The set of SNVs across the entire EBV genome is downsampled to 23, which is the same as the number of SNVs reported in the EBER region. NPC sample is analyzed in the test set using the same approach as the floating number of SNVs in the downsampled set of SNVs. In the downsampling analysis, a certain number (such as 23, 25, 100, 200, or 500, etc.) of SNVs are randomly selected from the 661 important SNVs. Then, for the test sample, the SNV sites within the downsampled set of SNVs covered by the EBV DNA sequence reads are identified. Then, in the training set of the covered and downsampled SNV sites, the NPC risk score training model is obtained by training the model using the genotype patterns of NPC samples and non - NPC samples. Through this training, the weight of the genotype at each site for the training model is determined. And by applying its own genotype pattern across these covered and downsampled SNV sites to the NPC risk score training model weighted by the downsampled SNV sites, the NPC risk score of the test sample is derived. The prediction performance of the NPC risk score training model with various numbers of SNV sites is summarized in Table 10. For a given number of SNV sites, SNVs are randomly selected for downsampling 10 times, and the area under the curve value in Table 10 is the average result of the 10 - time random downsampling. The discrimination between NPC and non-NPC samples was evaluated by ROC curve analysis. The area under the curve was 0.78. This value was higher than that (0.72) obtained by analyzing the genotype patterns of 23 SNVs reported in the EBER region.

Table 10

[0206] This study reports the analysis of EBV genotype information by plasma DNA sequencing. Through paired-end sequencing, the molecular characteristics of plasma EBV DNA molecules, including number and size, were identified between NPC subjects with plasma EBV DNA and non-NPC subjects. By incorporating such count- and size-based analysis of plasma EBV DNA, the positive predictive value of current PCR-based protocols can be almost doubled, which can form the basis for a second-generation sequencing-based screening test. Sequencing of plasma samples from NPC and non-NPC subjects can additionally obtain EBV genotype information and enhance its potential clinical utility. The NPC risk score can be used to be determined by viral genome-wide markers rather than a single gene marker. Here, the risk score was derived based on the mutation pattern through the identification of SNV sites across the EBV genome. Plasma sequencing of EBV genotype information may include sequencing of plasma samples with low-concentration EBV DNA molecules, and thus, the coverage of the EBV genome will result in incomplete results. In some cases, it is possible that no EBV DNA reads cover informative SNV sites.

[0207] NPC risk score can be used to be determined by viral genome-wide markers rather than a single gene marker. Here, the ri...

Claims

1. 1. A method of screening for a pathogen-associated disorder in a subject, comprising: determining a characteristic of cell-free nucleic acid molecules from a pathogen in a biological sample of said subject; receiving data from a first assay performed at a first time point, The characteristics of the cell-free nucleic acid molecules from the pathogen include quantity, methylation state, mutation pattern, fragment size, or cell-free nucleic acid from said subject in said biological sample a relative abundance of a gene encoding a pathogen-associated disorder in a subject compared to a gene encoding a pathogen-associated disorder in a subject; receiving an indication of a risk of developing harm; screening for said pathogen-associated disorder in said subject based on said characteristic. determining a second time point at which a second assay is performed to determine the identity of the first assay; determining that an interval between said first point and said second point in time is inversely correlated with said risk; A screening method comprising:

2. 1. A method for prognosing a pathogen-associated disorder in a subject, comprising: determining a characteristic of cell-free nucleic acid molecules from a pathogen in a biological sample of said subject; receiving data from a first assay, comprising: The characteristics of the cell-free nucleic acid molecules include amount, methylation state, mutation pattern, fragment size, or the relative presence in said biological sample compared to cell-free nucleic acid molecules from said subject. receiving, including a quantity; The characteristics of the cell-free nucleic acid molecule from the pathogen, as well as the age, the subject's smoking habits, the subject's family history of pathogen-associated disorders, the subject's genotypic factors, determining whether said subject is a dietary history or ... generating a report indicating a risk of developing said pathogen-associated disorder; A method for prognostic diagnosis comprising:

3. The results of the first assay result in medical treatment of the subject for the pathogen-associated disorder. The method of claim 1, wherein the method does not produce any results.

4. 3. The method of claim 2, wherein the medical treatment comprises therapeutic treatment, radiation therapy, or surgical treatment. The method described above.

5. The subject is assessed for the second time point by a clinical diagnostic test having a false positive rate of less than 1%.

5. The method of claim 1, 3 or 4, wherein the patient is diagnosed as not having the pathogen-associated disorder prior to the administration of the agent. The method described.

6. The clinical diagnostic test may be a physical examination, an invasive biopsy, an endoscopy, a magnetic resonance imaging, a positive emission tomography, or a positive emission tomography.

6. The method of claim 5, comprising laminography, computed tomography, or x-ray imaging.

7. The clinical diagnostic test is an invasive method including histological, cytological, or cellular nucleic acid analysis. The method of claim 5 , comprising a biopsy.

8. The interval is at least about 2 months, 4 months, 6 months, 8 months, 10 months, or 12 months. The method according to any one of claims 1, 3 or 7, wherein the

9. 9. The method of claim 8, wherein the interval is at least about 12 months.

10. The method of any one of claims 1 to 9, further comprising carrying out the first assay. Law.

11. conducting said first assay, (i) obtaining a first biological sample from the subject; (ii) extracting a first amount of cell-free nucleic acid from the pathogen in the first biological sample; Measuring the child; The method of claim 10, comprising:

12. The first amount is determined by isolating the pathogen from the first biological sample. The method of claim 11, comprising determining the copy number of the nucleic acid molecule.

13. 13. The method of claim 11 or 12, wherein the measurement comprises polymerase chain reaction (PCR). method.

14. The method of claim 11 or 12, wherein the measuring comprises quantitative PCR (qPCR).

15. The first amount is a fraction of the cell-free nucleic acid fragment from the pathogen in the first biological sample. The method of claim 11 , comprising measuring a first percentage of offspring.

16. The first assay comprises: (iii) obtaining a second biological sample from the subject if the first amount exceeds a threshold value. and obtaining a second amount of the pathogen in the second biological sample. The method of any one of claims 11 to 15, further comprising measuring cellular nucleic acid molecules.

17. The second biological sample is obtained about four weeks after the first biological sample. The method of claim 16 .

18. The interval between the first time point and the second time point is such that the second amount falls below the threshold. the first amount and the second copy number both exceed the threshold value, compared to an interval where the first amount and the second copy number are 18. The method according to claim 16 or 17, wherein the rotation is shorter.

19. The interval between the first time point and the second time point is such that the first amount exceeds the threshold. the interval during which the first amount is below the threshold is longer compared to the interval during which the first amount is below the threshold.

19. The method according to any one of 16 to 18.

20. The interval between the first time point and the second time point is a function of the first amount and the second amount.

20. The method according to claim 16, wherein the amount of the serotonin-containing compound is more than about 1 year. method.

21. The interval between the first time point and the second time point is such that the second amount falls below the threshold. The method according to any one of claims 16 to 20, wherein the period of time required for the treatment is about 2 years.

22. The interval between the first time point and the second time point is such that the first amount falls below the threshold. The method according to any one of claims 16 to 21, wherein the period of time required for the treatment is about 4 years.

23. The first assay comprises: determining the methylation state of the cell-free nucleic acid molecules from the pathogen in the biological sample; The method of claim 10, comprising determining

24. Determining the methylation state of the cell-free nucleic acid molecules in the biological sample 24. The method of claim 23, further comprising treating with a sensitive restriction enzyme or bisulfite. 。

25. The determining of the methylation status comprises determining the methylation status of cell-free nucleic acids in the biological sample of the subject.

24. The method of claim 23, further comprising performing nucleic acid recognition sequencing.

26. The methylation-aware sequencing is carried out by converting unmethylated cytosines to uracils.

26. The method of claim 25, comprising a phytoconversion.

27. 26. The method of claim 25, wherein the methylation-aware sequencing comprises treatment with a methylation-sensitive restriction enzyme. The method described above.

28. The first assay comprises: The pathogen in the biological sample is isolated from the cell-free nucleic acid molecule fragment size. The method of claim 10, further comprising determining a distribution of the noise.

29. Determining the fragment size distribution identifies the location of cell-free nucleic acid molecules in the biological sample. performing sequence determination and determining the sequence linkages mapped to the reference genome of the pathogen; and extracting fragments of the cell-free nucleic acid molecules from the pathogen in the biological sample based on the fragments. and determining a segment size.

30. The first assay comprises: determining a mutation pattern of the cell-free nucleic acid molecule from the pathogen in the biological sample; The method of claim 10, comprising determining

31. Determining the mutation pattern includes sequencing cell-free nucleic acid molecules in the biological sample. and performing a sequence analysis based on sequence reads mapped to the reference genome of the pathogen. and determining the mutation pattern of the cell-free nucleic acid molecule from the pathogen in the biological sample. and determining a time for which the first and second digits are to be processed.

32. The mutation pattern of the cell-free nucleic acid molecule from the pathogen comprises a single base mutation.

32. The method according to item 30 or 31.

33. Identifying the mutation pattern The sequence reads mapped to the reference genome of the pathogen, 33. The method of claim 32, comprising determining a level of similarity between the somatic disorder-associated reference genome.

34. The disorder-associated reference genome of the pathogen is a genome of the pathogen identified in diseased tissue.

34. The method of claim 33, comprising:

35. determining said level of similarity Separating the reference genome of the pathogen into a plurality of bins; A similarity index for each of the plurality of bins to the disorder-associated reference genome of the pathogen. determining a number of similarity indices that map to the reference genome of the pathogen; At least one of the sequence reads matched to the disorder-associated reference genome of the pathogen. This correlates with the proportion of mutation sites in each bin that have the same nucleotide variant as the corresponding bin.

35. The method of claim 33 or 34, wherein determining the level of similarity comprises:

36. The disorder-associated reference genome for the pathogen comprises a plurality of disorder-associated reference genomes for the pathogen. and said determining the similarity level comprises: For each of the plurality of disorder-associated reference genomes of the pathogen, determining a similarity index for each of them respectively; The plurality of disorders in which each of the similarity indices in each of the bins is greater than a cutoff value. determining a bin score for each of the plurality of bins based on a proportion of the associated reference genome; 36. The method of claim 35, comprising:

37. The lengths of the bins are about 100, 200, 300, 400, 500, 60 37. The method of claim 35 or 36, wherein the length is 0, 700, 800, 900, or 1000 bp. How to.

38. The first assay comprises detecting the amount of cell-free nucleic acid from the pathogen in the biological sample. the methylation status, the fragment size distribution, or the mutation pattern of the offspring The method of any of claims 10 to 37, comprising determining:

39. a data set comprising the characteristics of the cell-free nucleic acid molecules from the pathogen in the biological sample; A classifier applied to the data input is used to determine the risk of the subject developing the pathogen-associated disorder. and calculating a score for the classifier in the biological sample. applying a function to the data input including the characteristics of the cell-free nucleic acid molecules from the pathogen. and a risk profile for assessing the risk of the subject developing the disorder. The method according to any one of claims 1 to 38, further comprising: generating an output including a core. Method of posting.

40. 40. The method of claim 39, wherein the classifier is trained on a labeled data set. Method of posting.

41. 2. The method of claim 1, further comprising performing the second assay at the second time point. method.

42. 42. The method of claim 41, wherein the second assay is identical to the first assay.

43. The second assay comprises assaying cell-free nucleic acid molecules from the subject, assaying for an invasive disease in the subject, the subject's biopsy, endoscopy, or magnetic resonance imaging of the subject.

42. The method according to claim 41.

44. 1. A method for analyzing nucleic acid molecules from a biological sample from a subject, comprising: a computer system for extracting cell-free nucleic acid from the biological sample of the subject; obtaining molecular sequence reads, the biological sample being obtaining cell-free nucleic acid molecules from a subject and potentially from a pathogen; In the computer system, the sequence reads of the cell-free nucleic acid molecules are aligning the genome of the pathogen to a reference genome; In the computer system, a mutation pattern of the cell-free nucleic acid molecule from the pathogen is detected. Identifying turns, wherein the mutation pattern corresponds to the reference genome of the pathogen. At each of the plurality of mutation sites above, a previous sequence mapped to the reference genome of the pathogen is characterizing nucleotide variants in the sequence reads; and identifying the plurality of variant sites as a variant of the disease; comprising at least 30 sites across said reference genome of a substance; and a status that identifies, or is indicative of, a pathogen-associated disorder status or risk thereof in said subject; Top and Including, a method of analysis.

45. The plurality of mutation sites comprises at least 40 mutations across the reference genome of the pathogen; At least 50, at least 60, at least 70, at least 80, at least 90, At least 100, at least 200, at least 300, at least 400, at least At least 500, at least 600, at least 700, at least 800, at least 900 , including at least 1000, at least 1100, or at least 1200 sites 45. The method of claim 44.

46. The plurality of mutation sites is at least 600 across the reference genome of the pathogen.

45. The method of claim 44, comprising a site.

47. The plurality of mutation sites comprises approximately 660 sites across the reference genome of the pathogen.

45. The method of claim 44, comprising:

48. The plurality of mutation sites are at least 1000 across the reference genome of the pathogen. The method of claim 44, comprising the site:

49. The plurality of mutation sites comprises approximately 1100 sites across the reference genome of the pathogen.

45. The method of claim 44, comprising:

50. The sequence in which the plurality of mutation sites are mapped to the reference genome of the pathogen. All reads that have a nucleotide variation that differs from the reference genome of the pathogen. The method of claim 44, comprising the site:

51. The sequence reads are aligned to the reference genome of the pathogen. Between the sequence read and the reference genome of the pathogen, and wherein the nucleic acid sequence is adapted to tolerate a maximum mismatch of 1, 5, 4, 3, 2, or 1 base. The method according to any one of claims 44 to 50.

52. The sequence reads are aligned to the reference genome of the pathogen. A maximum mismatch of 2 bases between the sequence read and the reference genome of the pathogen. The method of any of claims 44 to 50, configured to allow

53. The mutation pattern of the sequence reads mapped to the reference genome of the pathogen. and diagnosing, prognosing, or monitoring the pathogen-associated disorder in the subject based on the results of the study. The method of any of claims 44 to 52, further comprising tarring.

54. The mutation pattern of the cell-free nucleic acid molecule from the pathogen comprises a single base mutation.

54. The method according to any one of items 44 to 53.

55. Identifying the mutation pattern The sequence reads mapped to the reference genome of the pathogen, 55. The method of claim 44, further comprising determining a level of similarity between the disorder-associated reference genome of the present invention and the target genome of the present invention. Any method as described above.

56. The disorder-associated reference genome of the pathogen is a genome of the pathogen identified in diseased tissue.

56. The method of claim 55, comprising:

57. determining said level of similarity Separating the reference genome of the pathogen into a plurality of bins; the similarity of each of the plurality of bins to the disorder-associated reference genome of the pathogen; determining a similarity index, the similarity index being a matched match to the reference genome of the pathogen; At least one of the sequence reads sequenced corresponds to the disorder-associated reference genome of the pathogen. The proportion of mutation sites in each bin that have the same nucleotide mutation as the nomenclature is correlated with the proportion of mutation sites in each bin that have the same nucleotide mutation as the nomenclature.

57. The method of claim 55 or 56, comprising:

58. The disorder-associated reference genome for the pathogen comprises a plurality of disorder-associated reference genomes for the pathogen. and said determining the similarity level comprises: For each of the plurality of disorder-associated reference genomes of the pathogen, determining a respective similarity index for each of the The plurality of disorders in which each of the similarity indices in each of the bins is greater than a cutoff value. determining a bin score for each of said plurality of bins based on the proportion of related reference genomes. and

59. 59. The method of claim 58, wherein the cutoff value is about 0.

9.

60. The lengths of the bins are about 100, 200, 300, 400, 500, 60 Any of claims 57 to 59, which is 0, 700, 800, 900, or 1000 bp. The method described above.

61. a data input including the mutation pattern of the cell-free nucleic acid molecule from the pathogen; The classifier is used to calculate a risk score for the subject developing an early pathogen-associated disorder. wherein the classifier determines the mutation pattern of the cell-free nucleic acid molecules from the pathogen. and configured to apply a function to a data input including a number of data points to determine whether the subject suffers from the disorder. generating an output including the risk score assessing the risk. The method according to any one of claims 44 to 60.

62. 62. The method of claim 61, wherein the classifier is trained on a labeled data set. Method of posting.

63. The classifier may be a Naive Bayes model, a logistic regression, a random forest, or a deterministic model. Decision Trees, Gradient Boosting Trees, Neural Networks, Deep Learning, Linear Shape / kernel support vector machine (SVM), linear / nonlinear regression, or linear discriminant analysis 63. The method of claim 61 or 62, comprising a mathematical model using analysis.

64. The method of any one of claims 44 to 63, wherein the pathogen is a virus.

65. 65. The method of claim 64, wherein the virus is Epstein-Barr virus (EBV). How to post

66. The pathogen-associated disorder is nasopharyngeal carcinoma, NK cell lymphoma, Burkitt's lymphoma, post-transplant lymphoma, 66. The method of claim 65, comprising a lymphoproliferative disease or Hodgkin's lymphoma.

67. The mutation pattern of the cell-free nucleic acid molecule from the pathogen is compared to the EBV reference genome (AJ At least one selected from the genomic sites listed in Table 6 associated with Both are 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 4 Each of said plurality of mutation sites includes 50, 500, 550, or 600 sites. The nucleotides of the sequence code mapped to the reference genome of the pathogen are 67. The method of claim 65 or 66, wherein the mutant is characterized.

68. Table 6, in which the multiple mutation sites are related to the EBV reference genome (AJ507799.2).

68. The method of claim 67, comprising a genomic site described in.

69. The mutation pattern of the cell-free nucleic acid molecule from the pathogen is compared to the EBV reference genome (AJ Randomly selected from the genomic sites listed in Table 6 related to the and mapping the plurality of mutation sites to the reference genome of the pathogen.

67. The method of claim 65 or 66, further comprising characterizing nucleotide variants of said sequence code. The method described.

70. The mutation pattern of the cell-free nucleic acid molecule from the pathogen is compared to the EBV reference genome (AJ Randomly selected from the genomic sites listed in Table 6 related to the At least 30, 40, 50, 100, 150, 200, 250, 300, 350 , 400, 450, 500, 550, or 600 sites. The sequence codes mapped to the reference genome of the pathogen in each of the 67. The method of claim 65 or 66, wherein the nucleotide variants are characterized.

71. 65. The method of claim 64, wherein the virus is a human papilloma virus (HPV). 。

72. 72. The method of claim 71, wherein the pathogen-associated disorder comprises cervical cancer, oropharyngeal cancer, or head and neck cancer. How to.

73. 65. The method of claim 64, wherein the virus is Hepatitis B virus (HBV).

74. 74. The method of claim 73, wherein the pathogen-associated disorder comprises cirrhosis or hepatocellular carcinoma (HCC). method.

75. The mutation pattern is indicative of a pathogen-associated disorder status in the subject, and The condition of a pathogenicity-associated disorder may be characterized by the presence of the pathogenicity-associated disorder in the subject, the amount of tumor tissue in said subject, the size of tumor tissue in said subject, the tumor in said subject the stage of the disease in the subject, the tumor burden in the subject, or the presence of tumor metastases in the subject The method according to any one of claims 44 to 74.

76. The biological sample may be whole blood, plasma, serum, urine, cerebrospinal fluid, buffy coat, vaginal fluid, Vaginal wash, saliva, oral rinse, nasal wash, nasal brush samples, and combinations thereof The method of any one of claims 44 to 75, wherein the compound is selected from the group consisting of a combination of

77. Execution by one or more computer processors of any of claims 1 to 76. A non-transitory computer readable medium containing machine executable code implementing a method.

78. Controlling a computer system to perform the operations of the method of any one of claims 1 to 76. A computer including a non-transitory computer-readable medium storing a plurality of instructions for carrying out Data products.

79. 79. A computer product according to claim 78; one or more processors for executing instructions stored on the computer-readable medium; , a system including

Citation Information

Patent Citations

  • Methods and systems for tumor detection

    WO2018081130A1

  • Cancer detection and classification using methylome analysis

    WO2019010564A1

Cited By

  • Method for early non-invasive multi-wavelength express diagnostics of polypous rhinosinusitis using raman scattering

    RU2862795C1