Detection of contaminated samples for cancer classification

JP2024544597A5Pending Publication Date: 2025-11-10GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024530567
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2022-11-23
Publication Date
2025-11-10

AI Technical Summary

Technical Problem

Current methods for DNA methylation profiling in cancer detection are hindered by the need for improved techniques to detect contamination of non-individual cell-free DNA (cfDNA) fragments, which can affect the accuracy of cancer classification.

Method used

A system and method for identifying and estimating contamination in cfDNA samples by using homozygous haplotype-based contamination markers, such as SNP and indel sites, to improve cancer classification accuracy by filtering out contaminated fragments.

Benefits of technology

Enhances the reliability of cancer classification by accurately detecting and removing contaminated cfDNA fragments, thereby improving the precision of disease prediction and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000061_0000
    Figure 00000061_0000
  • Figure 00000061_0001
    Figure 00000061_0001
  • Figure 00000061_0002
    Figure 00000061_0002
Patent Text Reader

Abstract

A method and system for detecting contaminant fragments in biological samples for cancer classification is disclosed. The system identifies multiple SNP site contaminant markers and indel site contaminant markers. The multiple SNP site contaminant markers include at least two SNP sites within a threshold distance, have population haplotype frequencies within a threshold frequency range, exclude guanine-adenine polymorphisms and / or cytosine-thymine polymorphisms, ensure Hardy-Weinberg equilibrium, or include any combination of the above parameters. The indel site contaminant markers include indel sequences that are within a threshold length, have high complexity, have population haplotype frequencies within a threshold frequency range, ensure Hardy-Weinberg equilibrium, or include any combination of the above parameters. The system identifies contaminant markers for which the sample is homozygous. The system estimates the contamination level of the sample by identifying fragments that have haplotypes that are different from the homozygous haplotype of the respective contaminant marker site.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 282,509, filed November 23, 2021, which is incorporated by reference in its entirety.

[0002] The present disclosure relates to detection of contaminating samples for cancer classification. [Background technology]

[0003] Methylation of deoxyribonucleic acid (DNA) plays an important role in the control of gene expression. Aberrant DNA methylation is involved in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS)) is increasingly recognized as a valuable diagnostic tool for cancer detection, diagnosis, and / or cancer monitoring. For example, specific patterns of differentially methylated regions and / or allele-specific methylation patterns are useful as molecular markers for non-invasive diagnosis using circulating cell-free (cf)DNA. However, there remains a need in the art for improved methods for detecting contamination of cfDNA fragments that are not derived from the individual of the sample.

[0004] The present disclosure is directed to addressing the above-mentioned problems. The background discussion provided herein is intended to generally present the context of the present disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims of this application, and no admission of prior art or suggestion of prior art is made by inclusion in this section. Summary of the Invention [Problem to be solved by the invention]

[0005] Early detection of disease conditions (such as cancer) in subjects is important because it allows for early treatment and therefore increases the chances of survival. Sequencing of DNA fragments in cell-free (cf)DNA samples can be used to identify features that can be used to classify disease. For example, in the assessment of cancer, features based on cell-free DNA from blood samples (such as the presence or absence of somatic mutations, methylation status, or other genetic abnormalities) can provide insight into whether a subject may have cancer, and further insight into what type of cancer the subject may have. To this end, included herein are systems and methods for analyzing cell-free DNA (cfDNA) sequencing data to determine the likelihood that a subject has a disease.

[0006] The present disclosure addresses the above-identified problems by providing an improved system and method for sample contamination detection of contaminant fragments for cancer classification. The system identifies one or more contaminant markers from a plurality of contaminant markers for which the sample has a homozygous haplotype. The system identifies any cfDNA fragments in the sample that have a haplotype at one of the identified contaminant markers that is different from the homozygous haplotype of the respective contaminant marker as a contaminated cfDNA fragment. The system estimates the contamination level of the sample based on the contaminant cfDNA fragments. The system can perform this contamination detection on both the training samples used to train the cancer classifier and also on the test samples in deploying the cancer classifier. The contaminant markers include a plurality of single nucleotide polymorphism (SNP) site contaminant markers and / or indel site contaminant markers. In some examples, the multiple SNP site contaminating marker comprises at least two SNP sites within a threshold distance, having a population haplotype frequency within a threshold frequency range, excluding guanine-adenine polymorphisms and / or cytosine-thymine polymorphisms, ensuring Hardy-Weinberg equilibrium, or any combination of the above parameters. In some examples, the indel site contaminating marker comprises an indel sequence within a threshold length, having high complexity, having a population haplotype frequency within a threshold frequency range, ensuring Hardy-Weinberg equilibrium, or any combination of the above parameters. [Means for solving the problem]

[0007] According to a first aspect, a method of predicting the presence of cancer in a test sample is disclosed, comprising the steps of obtaining the test sample comprising a plurality of sequence reads for cell-free DNA (cfDNA) fragments in the test sample; identifying one or more contaminant markers from a plurality of contaminant markers for which the test sample has a homozygous haplotype; for each of the identified one or more contaminant markers for which the test sample has a homozygous haplotype, identifying as contaminant cfDNA fragments any cfDNA fragments in the test sample having a haplotype at one of the identified contaminant markers that is different from the homozygous haplotype of the respective contaminant marker; estimating a contamination level based on any of the identified contaminant cfDNA fragments; and determining whether the contamination level is below a threshold level; and in response to determining that the contamination level is below the threshold level, performing a cancer classification on the sequence reads of the cfDNA fragments in the test sample to generate a cancer prediction.

[0008] In the method of the first aspect, the plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

[0009] In the method of the first aspect, the plurality of SNP sites are within 10 base pairs (bps).

[0010] In the method of the first aspect, said plurality of SNP sites have a population haplotype frequency within the range of 45% to 55%.

[0011] In the method of the first aspect, said plurality of SNP sites excludes guanine-adenine and cytosine-thymine polymorphisms.

[0012] In the method of the first aspect, said haplotypes at each of the plurality of SNP sites are in Hardy-Weinberg equilibrium.

[0013] In the method of the first aspect, said plurality of contaminant markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 SNP sites of said plurality.

[0014] In the method of the first aspect, said plurality of contaminant markers comprises said plurality of SNP sites in Table 1.

[0015] In the method of the first embodiment, the plurality of contaminant markers comprises insertion-deletion (indel) sites.

[0016] In the method of the first aspect, the indel site is 5 bps to 10 bps.

[0017] In the method of the first aspect, the indel site has a population haplotype frequency within the range of 45% to 55%.

[0018] In the method of the first embodiment, said haplotypes at each indel site are in Hardy-Weinberg equilibrium.

[0019] In the method of the first embodiment, said plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of said indel sites.

[0020] In the method of the first aspect, the plurality of contamination markers comprises the indel sites of Table 3.

[0021] In the method of the first embodiment, each contaminant marker comprises a probe designed to target each haplotype of said contaminant marker.

[0022] In the method of the first aspect, the step of estimating the contamination level is further performed based on one or more of the number of identified contaminating cfDNA fragments, the sequencing depth of the test sample, the number of the cfDNA fragments in the test sample, and the number of the contaminating markers.

[0023] In the method of the first aspect, in response to determining that said contamination level is equal to or greater than said threshold level, classification of cancer is withheld.

[0024] In the method of the first aspect, said cancer prediction comprises a binary prediction between cancer and non-cancer.

[0025] In the method of the first aspect, the cancer prediction comprises a multi-class cancer prediction among multiple cancer types.

[0026] In the method of the first aspect, the step of performing the cancer classification further comprises a step of filtering the initial set of cfDNA fragments of the test sample with p-value filtering to generate a set of abnormal fragments, the filtering removing fragments from the initial set that are below a threshold p-value with respect to other fragments to generate the set of abnormal fragments.

[0027] In the method of the first aspect, the classification model is a machine learning model.

[0028] According to a second aspect, a non-transitory computer readable storage medium is disclosed storing instructions which, when executed by a computer processor, cause the computer processor to perform the method of the first aspect.

[0029] According to a third aspect, a system is disclosed comprising a computer processor and the non-transitory computer readable storage medium of the second aspect.

[0030] According to a fourth aspect, a method for predicting the presence of a disease in a test sample is disclosed, comprising the steps of obtaining a test sample comprising a plurality of sequence reads for cell-free DNA (cfDNA) fragments in the test sample; identifying one or more contaminant markers from a plurality of contaminant markers for which the test sample has a homozygous haplotype; for each of the one or more contaminant markers for which the test sample has a homozygous haplotype, identifying as a contaminant cfDNA fragment any of the cfDNA fragments in the test sample having a haplotype at one of the identified contaminant markers that is different from the homozygous haplotype of the respective contaminant marker; estimating a contamination level based on the identified contaminant cfDNA fragments; and determining whether the contamination level is below a threshold level; and in response to determining that the contamination level is below the threshold level, performing a disease classification on the sequence reads of the cfDNA fragments in the test sample to generate a disease prediction.

[0031] In the method of the fourth aspect, said plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

[0032] In the method of the fourth aspect, the plurality of SNP sites are within 10 base pairs (bps).

[0033] In the method of the fourth aspect, said plurality of SNP sites have a population haplotype frequency within the range of 45% to 55%.

[0034] In the method of the fourth embodiment, said plurality of SNP sites excludes guanine-adenine and cytosine-thymine polymorphisms.

[0035] In the method of the fourth embodiment, said haplotypes at each of the multiple SNP sites are in Hardy-Weinberg equilibrium.

[0036] In the method of the fourth aspect, said plurality of contaminant markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 SNP sites of said plurality.

[0037] In the method of the fourth aspect, said plurality of contaminant markers comprises said plurality of SNP sites in Table 1.

[0038] In the method of the fourth embodiment, the plurality of contaminant markers comprises insertion-deletion (indel) sites.

[0039] In the method of the fourth aspect, the indel site is 5 bps to 10 bps.

[0040] In the method of the fourth aspect, the indel site has a population haplotype frequency within the range of 45% to 55%.

[0041] In the method of the fourth embodiment, said haplotypes at each indel site are in Hardy-Weinberg equilibrium.

[0042] In the method of the fourth embodiment, said plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of said indel sites.

[0043] In the method of the fourth aspect, said plurality of contamination markers comprises said indel sites of Table 3.

[0044] In the method of the fourth embodiment, each contaminant marker comprises a probe designed to target each haplotype of said contaminant marker.

[0045] In the method of the fourth aspect, the step of estimating the contamination level is further performed based on one or more of the number of identified contaminating cfDNA fragments, the sequencing depth of the test sample, the number of cfDNA fragments in the test sample, and the number of contaminating markers.

[0046] In the method of the fourth aspect, in response to determining that said contamination level is equal to or greater than said threshold level, disease classification is withheld.

[0047] In the method of the fourth aspect, said disease prediction comprises a binary prediction between disease and disease-free.

[0048] In the method of the fourth aspect, the disease prediction comprises a multi-class cancer prediction among multiple diseases.

[0049] In the method of the fourth aspect, the step of performing the disease classification includes the steps of generating a test feature vector based on the sequence reads of the cfDNA fragments in the test sample, and inputting the test feature vector into a classification model to generate the disease prediction for the test sample.

[0050] In the method of the fourth aspect, the step of performing the disease classification further comprises filtering the initial set of cfDNA fragments of the test sample with p-value filtering to generate a set of abnormal fragments, the filtering step comprising removing fragments from the initial set that are below a threshold p-value with respect to other fragments to generate the set of abnormal fragments, and the test feature vector is based on the sequence reads of the set of abnormal fragments.

[0051] In the method of the fourth aspect, the classification model is a machine learning model.

[0052] According to a fifth aspect, a non-transitory computer readable storage medium is disclosed having stored thereon instructions that, when executed by a computer processor, cause the computer processor to perform the method of the fourth aspect.

[0053] According to a sixth aspect, a system is disclosed comprising a computer processor and the non-transitory computer readable storage medium of the fifth aspect.

[0054] According to a seventh aspect, a method for predicting the presence of contamination in a test sample is disclosed, comprising the steps of obtaining sequence reads derived from a plurality of cell-free DNA (cfDNA) fragments in the test sample; identifying one or more contaminant markers from a plurality of contaminant markers for which the test sample has a homozygous haplotype based on the sequence reads; identifying any cfDNA fragments in the test sample having a haplotype at one of the identified contaminant markers different from the homozygous haplotype of each of the contaminant markers as contaminant cfDNA fragments; estimating a contamination level based on any of the identified contaminant cfDNA fragments; and determining whether the contamination level is below a threshold level; and in response to determining that the contamination level is below the threshold level, generating a notification indicating that the test sample is contaminated.

[0055] In the method of the seventh aspect, said plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

[0056] In the method of the seventh aspect, the plurality of SNP sites are within 10 base pairs (bps).

[0057] In the method of the seventh aspect, said plurality of SNP sites have a population haplotype frequency within the range of 45% to 55%.

[0058] In the method of the seventh embodiment, said plurality of SNP sites excludes guanine-adenine and cytosine-thymine polymorphisms.

[0059] In the method of the seventh embodiment, said haplotypes at each of the multiple SNP sites are in Hardy-Weinberg equilibrium.

[0060] In the method of the seventh aspect, said plurality of contaminant markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 SNP sites of said plurality.

[0061] In the method of the seventh embodiment, said plurality of contaminant markers comprises said plurality of SNP sites from Table 1.

[0062] In the method of the seventh embodiment, the plurality of contaminant markers comprises insertion-deletion (indel) sites.

[0063] In the method of the seventh aspect, the indel site is 5 bps to 10 bps.

[0064] In the method of the seventh aspect, the indel site has a population haplotype frequency in the range of 45% to 55%.

[0065] In the method of the seventh embodiment, said haplotypes at each indel site are in Hardy-Weinberg equilibrium.

[0066] In the method of the seventh aspect, said plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of said indel sites.

[0067] In the method of the seventh embodiment, said plurality of contamination markers comprises said indel sites in Table 3.

[0068] In the method of the seventh embodiment, each contaminant marker comprises a probe designed to target each haplotype of said contaminant marker.

[0069] In the method of the seventh aspect, the step of estimating the contamination level is further performed based on one or more of the number of identified contaminating cfDNA fragments, the sequencing depth of the test sample, the number of the cfDNA fragments in the test sample, and the number of the contaminating markers.

[0070] According to an eighth aspect, a non-transitory computer readable storage medium is disclosed having stored thereon instructions which, when executed by a computer processor, cause the computer processor to perform the method of the seventh aspect.

[0071] According to a ninth aspect, a system is disclosed comprising a computer processor and the non-transitory computer readable storage medium of the eighth aspect.

[0072] According to a tenth aspect, a method for training a cancer classification model is disclosed, comprising the steps of: obtaining a plurality of training samples, including a first training sample, each training sample comprising a plurality of cell-free DNA (cfDNA) fragments; obtaining, for each training sample, sequence reads derived from the cfDNA fragments in the training sample; identifying, for the first training sample, one or more contaminant markers from a plurality of contaminant markers for which the first training sample has a homozygous haplotype based on the sequence reads of the first training sample; and determining, for the first training sample, the homozygosity of each of the contaminant markers at one of the identified contaminant markers. identifying any cfDNA fragments in the first training sample having a haplotype different from a haplotype as a contaminant cfDNA fragment; estimating a contamination level for the first training sample based on any identified contaminant cfDNA fragments and determining for the first training sample whether the contamination level is below a threshold level; and removing the first training sample from the plurality of training samples in response to determining that the contamination level of the first training sample is above the threshold level, wherein the plurality of training samples excluding the first training sample is used to train the cancer classification model to generate a cancer prediction for the test sample.

[0073] In the method of the tenth aspect, the plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

[0074] In the method of the tenth aspect, the multiple SNP sites are within 10 base pairs (bps).

[0075] In the method of the tenth aspect, said plurality of SNP sites has a population haplotype frequency within the range of 45% to 55%.

[0076] In the method of the tenth aspect, the plurality of SNP sites excludes guanine-adenine and cytosine-thymine polymorphisms.

[0077] In the method of the tenth aspect, said haplotypes at each of the multiple SNP sites are in Hardy-Weinberg equilibrium.

[0078] In the method of the tenth aspect, said plurality of contaminant markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 SNP sites of said plurality.

[0079] In the method of the tenth aspect, the plurality of contaminant markers comprises the plurality of SNP sites in Table 1.

[0080] In the method of the tenth aspect, the plurality of contaminant markers comprises insertion-deletion (indel) sites.

[0081] In the method of the tenth aspect, the indel site is 5 bps to 10 bps.

[0082] In the method of the tenth aspect, the indel site has a population haplotype frequency in the range of 45% to 55%.

[0083] In the method of the tenth embodiment, the haplotypes at each indel site are in Hardy-Weinberg equilibrium.

[0084] In the method of the tenth aspect, said plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of said indel sites.

[0085] In the method of the tenth aspect, the plurality of contamination markers comprises the indel sites of Table 3.

[0086] In the method of the tenth embodiment, each contaminant marker comprises a probe designed to target each haplotype of said contaminant marker.

[0087] In the method of the tenth aspect, the plurality of training samples comprises a first cohort of non-cancerous samples and a second cohort of cancer samples, and the cancer classification model is trained to determine the likelihood of the presence of cancer.

[0088] In the method of the tenth embodiment, the second cohort of cancer samples comprises one or more samples having a first cancer type and one or more additional samples having a second cancer type, wherein the cancer classification model is trained to determine a first likelihood of the presence of the first cancer type and a second likelihood of the presence of the second cancer type.

[0089] In the method of the tenth aspect, said cancer classification model is a machine learning model.

[0090] In the method of the tenth embodiment, the cancer classification model is at least one of a decision tree, a neural network, a multi-layer perceptron, and a support vector machine.

[0091] According to an eleventh aspect, a non-transitory computer readable storage medium is disclosed having stored thereon instructions which, when executed by a computer processor, cause the computer processor to perform the method of the tenth aspect.

[0092] According to a twelfth aspect, a system is disclosed comprising a computer processor and the non-transitory computer readable storage medium of the eleventh aspect.

[0093] According to a thirteenth aspect, a computer program product is disclosed comprising a non-transitory computer readable storage medium, the computer program product comprising the non-transitory computer readable storage medium storing the trained cancer classification model, the computer program product being produced by the method of the eleventh aspect.

[0094] According to a fourteenth aspect, a treatment kit is disclosed comprising one or more collection containers for storing a biological sample comprising genetic material from an individual, and a plurality of probes targeting a plurality of contamination markers, the plurality of probes comprising at least one of Tables 2 and 4.

[0095] In the treatment kit of the fourteenth aspect, the plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

[0096] In the treatment kit of the fourteenth aspect, the plurality of SNP sites are within 10 base pairs (bps).

[0097] In the treatment kit of the fourteenth aspect, the plurality of SNP sites have a population haplotype frequency within the range of 45% to 55%.

[0098] In the treatment kit of the fourteenth embodiment, the plurality of SNP sites excludes a guanine-adenine polymorphism and a cytosine-thymine polymorphism.

[0099] In the treatment kit of the fourteenth aspect, said haplotypes at each of the multiple SNP sites are in Hardy-Weinberg equilibrium.

[0100] In the treatment kit of the fourteenth aspect, said plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 SNP sites of said plurality.

[0101] In the treatment kit of the fourteenth aspect, the plurality of contamination markers comprises the plurality of SNP sites of Table 1.

[0102] In the therapeutic kit of the fourteenth embodiment, the plurality of contamination markers comprises insertion-deletion (indel) sites.

[0103] In the treatment kit of the fourteenth aspect, the indel site is 5 bps to 10 bps.

[0104] In the treatment kit of the fourteenth aspect, the indel site has a population haplotype frequency within the range of 45% to 55%.

[0105] In the therapeutic kit of the fourteenth aspect, the haplotypes at each indel site are in Hardy-Weinberg equilibrium.

[0106] In the treatment kit of the fourteenth embodiment, the plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 indel sites.

[0107] In the treatment kit of the fourteenth aspect, the plurality of contamination markers comprises the indel sites of Table 3.

[0108] In the therapeutic kit of the fourteenth embodiment, each contamination marker comprises a probe designed to target each haplotype of said contamination marker.

[0109] The therapeutic kit of the fourteenth embodiment further comprises one or more reagents for isolating nucleic acid fragments in a biological sample.

[0110] The treatment kit of the fourteenth aspect further comprises the first computer program product including one or more of the non-transient computer readable storage medium of the second aspect, the non-transient computer readable storage medium of the fifth aspect, the non-transient computer readable storage medium of the eighth aspect, and the non-transient computer readable storage medium of the eleventh aspect.

[0111] The treatment kit of the fourteenth aspect further comprises the computer program product of the thirteenth aspect. [Brief description of the drawings]

[0112] [Figure 1] 1 is an exemplary flow chart illustrating a process of detecting contamination in a sample, according to one or more embodiments. [Figure 2A] 1 is an exemplary flow chart illustrating a process for identifying multiple SNP sites for use as contamination markers in contamination detection, according to one or more embodiments. [Figure 2B] 1 is an exemplary flowchart illustrating a process for identifying indel sites for use as contamination markers in contamination detection, according to one or more embodiments. [Figure 3A] 1 is an exemplary flow chart illustrating the process of sequencing fragments of cell-free (cf) DNA to obtain a methylation status vector, according to one or more embodiments. [Figure 3B] FIG. 3B is an exemplary illustration of the process of FIG. 3A of sequencing fragments of cell-free (cf) DNA to obtain a methylation status vector, according to one or more embodiments. [Figure 4A] 1 is a flowchart illustrating a process for generating a data structure for a healthy control group in accordance with one or more embodiments. [Figure 4B] 1 is an exemplary flow chart illustrating a process for identifying aberrantly methylated fragments from a sample, according to one or more embodiments. [Figure 5A]1 is an exemplary flowchart illustrating a process of training a cancer classifier, according to one or more embodiments. [Figure 5B] FIG. 1 illustrates an example of generating a feature vector used to train a cancer classifier, according to one or more embodiments. [Figure 6A] 1 shows an illustrative flow chart of an apparatus for sequencing a nucleic acid sample, according to one or more embodiments. [Figure 6B] FIG. 1 illustrates an example block diagram of an analysis system in accordance with one or more embodiments. [Figure 7] FIG. 13 shows the distribution of the ratio of the number of fragments called as contaminant to the total number of fragments called as reference or alternative for each contaminant marker called as homozygous for a given sample. [Figure 8] 1 shows a scatter plot of allele frequencies of contaminant markers according to a first group of example results. [Figure 9] FIG. 1 shows two graphs showing the genotype and zygosity of contaminant markers according to the first group of example results. [Figure 10] FIG. 11 is a graph showing the percentage of contaminating markers that are homozygous and overlap sufficient fragments for a given sample to be considered useful for estimating contamination of that sample, according to the exemplary results of the first group. [Figure 11A] 1 shows a graph of estimated contamination levels for samples of different batches according to a first example result. [Figure 11B] 13 shows additional graphs of estimated contamination levels for different batches of samples according to the first example result. [Figure 12] FIG. 1 shows the distribution of the number of unique cfDNA fragments obtained for each contaminant marker listed in Tables 2 and 4 according to the second set of example results. [Figure 13] FIG. 13 shows the results of the quality check process according to a second set of example results. [Figure 14] 2 shows a scatter plot of allele frequencies by second result set. [Figure 15]A scatter plot of the frequency with which markers are called heterozygous. [Figure 16] A scatter plot of the Hardy-Weinberg term values ​​for the second group of example results is shown below. [Figure 17] FIG. 13 shows the distribution of the percentage of the number of fragments classified as contaminant relative to the total number of fragments called reference or alternative for each contaminant marker called homozygous for a given sample according to the example results of the second group. [Figure 18] FIG. 13 illustrates the application of the formula for estimating contamination rates for a hypothetical set of fragments according to the example results of the second group. [Figure 19] FIG. 13 shows the results of applying the contamination fraction model to simulated data according to a second group of example results. [Figure 20] FIG. 13 shows the results of applying the contaminant fraction model to four titration pairs of cfDNA samples with varying titration levels according to the second set of example results. [Figure 21] FIG. 1 shows the distribution of estimated contamination fractions for the 84 cfDNA samples in the experiment according to the second example result. [Figure 22] FIG. 1 shows that for the 84 samples in the experiment, the estimates obtained when considering multiple SNP and indel markers separately on their own are highly correlated and do not show significant systematic bias, in accordance with the exemplary results of the second group. [Figure 23] FIG. 1 shows Table 1, which includes multiple SNP contaminant markers, according to one or more embodiments. [Figure 24] 2 shows Table 2, which includes a probe sequence list for multiple SNP contaminant markers, according to one or more embodiments. [Diagram 25] 1 shows Table 3, which includes indel contamination markers, according to one or more embodiments. [Figure 26] 1 shows Table 4, which includes a list of probe sequences for indel-contaminating markers, according to one or more embodiments.

[0113] The figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following considerations that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

Best Mode for Carrying Out the Invention

[0114] [I. Overview] <I.A. Overview of Methylation> In accordance with the present specification, cfDNA fragments from an individual are processed by methods such as converting unmethylated cytosine to uracil, the nucleotide sequence is determined, and the sequence reads are compared to a reference genome to identify the methylation status of specific CpG sites within the DNA fragment. Each CpG site may be either methylated or unmethylated. By identifying abnormal methylation fragments compared to healthy individuals, the cancer status of the subject can be determined. As is well known in the art, abnormal DNA methylation causes various effects (compared to healthy controls), which may contribute to cancer. Identifying abnormally methylated cfDNA fragments presents various challenges. First, determining that a DNA fragment is abnormally methylated has important implications when compared to a control group, but if the number of control groups is small, statistical variations occur due to the small size of the control group, and the determination loses reliability. Furthermore, even within the control group, the methylation status varies, and it is difficult to consider this when determining that a subject's DNA fragment is abnormally methylated. In addition, the methylation of cytosine at a certain CpG site may affect the methylation of subsequent CpG sites. Encapsulating this dependency relationship itself can be another challenge.

[0115] Methylation can typically occur in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation can occur in dinucleotides of cytosine and guanine, which are referred to herein as "CpG sites". In other examples, methylation can occur in cytosines that are not part of a CpG site, or in other nucleotides that are not cytosine, but these are rare phenomena. In the present disclosure, for the sake of clarity, methylation is discussed in relation to CpG sites. Aberrant DNA methylation can be identified as hypermethylation or hypomethylation, both of which may indicate a cancerous state. Throughout the present disclosure, if a DNA fragment contains CpG sites above a threshold and a threshold proportion of those CpG sites are methylated or unmethylated, hypermethylation and hypomethylation of the DNA fragment can be characterized.

[0116] The principles described herein can be similarly applied to the detection of methylation in non-CpG contexts, including non-cytosine methylation. In such embodiments, the wet laboratory assays used to detect methylation may be different from those described herein. Further, the methylation state vectors discussed herein can generally include elements that are sites where methylation has occurred or not occurred, even if those sites are not particularly CpG sites. With this substitution, the remaining processes described herein can be the same, and as a result, the concepts of the invention described herein can be applied to those other forms of methylation.

[0117] <I.B. Definitions> The term "cell free nucleic acid" or "cfNA" refers to nucleic acid fragments circulating in an individual's body (e.g., blood) that originate from one or more healthy cells and / or from one or more unhealthy cells (e.g., cancer cells). The term "cell free DNA" or "cfDNA" refers to deoxyribonucleic acid fragments circulating in an individual's body (e.g., blood). Additionally, cfNA or cfDNA in an individual's body may be derived from other sources other than human.

[0118] The terms "genomic nucleic acid," "genomic DNA," or "gDNA" refer to a nucleic acid or deoxyribonucleic acid molecule obtained from one or more cells. In various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or tumor cells (e.g., biopsy samples). In some embodiments, gDNA can be extracted from cells from the blood lineage, such as white blood cells.

[0119] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from tumor cells or other types of cancer cells and that may be released into an individual's bodily fluids (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dying cells, or actively released by viable tumor cells.

[0120] The terms "DNA fragment," "fragment," or "DNA molecule" may generally refer to any deoxyribonucleic acid fragment, i.e., cfDNA, gDNA, ctDNA, etc.

[0121] The terms "anomalous fragment," "anomalously methylated fragment," or "fragment with an anomalous methylation pattern" refer to a fragment that has aberrant methylation of CpG sites. The aberrant methylation of a fragment can be determined using a probability model to identify the surprise of observing the methylation pattern of the fragment in a control group.

[0122] The term "unusual fragment with extreme methylation" or "UFXM" refers to a hypomethylated or hypermethylated fragment, which has at least some (e.g., 5) CpG sites that are methylated or unmethylated above a threshold (e.g., 90%), respectively.

[0123] The term "anomaly score" refers to the score of a CpG site based on the number of aberrant fragments (or, in some embodiments, UFXMs) from a sample that overlap that CpG site. The anomaly score is used in the context of characterizing a sample for classification.

[0124] As used herein, the term "about" or "approximately" may mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which may depend in part on the way in which the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" may mean within one standard deviation or more than one standard deviation, according to the practice of one of ordinary skill in the art. "About" may mean within a range of ±20%, ±10%, ±5%, or ±1% of a given value. The term "about" or "approximately" may mean within an order of magnitude, within 5-fold, or within 2-fold of a value. When a particular value is described in the present specification and claims, the term "about" should be assumed to mean within an acceptable error range of the particular value, unless otherwise indicated. The term "about" may have the meaning commonly understood by one of ordinary skill in the art. The term "about" may refer to ±10%. The term "about" may refer to ±5%.

[0125] As used herein, the terms "biological sample", "patient sample", or "sample" refer to any sample taken from a subject that can reflect a biological state associated with the subject and that contains cell-free DNA. Examples of biological samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of a subject. A biological sample can include any tissue or material from a subject, living or dead. A biological sample can be a cell-free sample. A biological sample can include nucleic acid (e.g., DNA or RNA) or fragments thereof. The term "nucleic acid" can refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or hybrids or fragments thereof. The nucleic acid in the sample can be cell-free nucleic acid. A sample can be a liquid sample or a solid sample (e.g., a cell sample or a tissue sample). The biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from hydrocele (e.g., of the testes), vaginal washings, pleural fluid, peritoneal fluid, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body (e.g., thyroid, breast), etc. The biological sample can also be a stool sample. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free). The biological sample can be treated (e.g., centrifugation and / or cell lysis) to physically disrupt the structure of tissues or cells, thereby releasing intracellular components into solution. The solutions may further comprise enzymes, buffers, salts, detergents, etc., which may be used to prepare the sample for analysis.

[0126] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" refer to a sample from a subject who does not have a particular condition or is otherwise healthy. In one example, a method as disclosed herein can be performed on a subject with a tumor, where the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from or from a database. The reference can be, for example, a reference genome used to map the nucleic acid fragment sequences obtained from sequencing a sample from the subject. The reference genome can refer to a haploid or diploid genome, to which the nucleic acid fragment sequences from the biological sample and the constitutional sample can be aligned and compared. An example of a constitutional sample can include DNA of white blood cells obtained from a subject. In the case of a haploid genome, there is only one nucleotide at each locus. In the case of a diploid genome, heterozygous loci can be identified. Each heterozygous locus can have two alleles, and either allele can match up with alignment to that locus.

[0127] As used herein, the term "cancer" or "tumor" refers to an abnormal mass of tissue whose growth exceeds and is uncoordinated with the growth of normal tissue.

[0128] As used herein, the phrase "healthy" refers to a subject having good health. A healthy subject can exhibit an absence of malignant or non-malignant disease. A "healthy individual" can have other diseases or conditions unrelated to the condition being assayed and that are not normally considered "healthy."

[0129] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation tends to occur at cytosine and guanine dinucleotides, referred to herein as "CpG sites." Methylation can also occur at cytosines or nucleotides other than cytosine at CpG sites, but this is rare. Methylation abnormalities in cfDNA can be identified as hypermethylation or hypomethylation, both of which may indicate a cancer state. DNA methylation abnormalities can cause various effects (compared to healthy controls) and may contribute to cancer. The principles described herein are equally applicable to the detection of methylation in CpG contexts as they are to the detection of methylation in non-CpG contexts, including non-cytosine methylation. Additionally, methylation status vectors can contain elements that are vectors of methylated or non-methylated sites in general (even if those sites are not specifically CpG sites).

[0130] As used interchangeably herein, the term "methylation fragment" or "nucleic acid methylation fragment" refers to a sequence of the methylation status of each CpG site in a plurality of CpG sites, determined by methylation sequencing of a nucleic acid (e.g., a nucleic acid molecule and / or a nucleic acid fragment). In a methylation fragment, the position and methylation status of each CpG site in a nucleic acid fragment are determined based on alignment of sequence reads (e.g., obtained from sequencing of a nucleic acid) to a reference genome. A nucleic acid methylation fragment consists of the methylation status (e.g., a methylation status vector) of each CpG site of a plurality of CpG sites, and specifies the position of the nucleic acid fragment in the reference genome (e.g., designated by the position of the first CpG site in the nucleic acid fragment using a CpG index, or another similar index) and the number of CpG sites in the nucleic acid fragment. Alignment of sequence reads to a reference genome based on methylation sequencing of a nucleic acid molecule can be performed using a CpG index. As used herein, the term "CpG index" refers to a list of each of a plurality of CpG sites (e.g., CpG 1, CpG 2, CpG 3, etc.) in a reference genome, such as a human reference genome, which may be in an electronic format. The CpG index further includes, for each CpG site in the CpG index, a corresponding genomic position in the corresponding reference genome. In this way, each CpG site in each nucleic acid methylation fragment is indexed to a specific position in each reference genome, which can be determined using the CpG index.

[0131] As used herein, the term "true positive" (TP) refers to a subject having a condition. A "true positive" can refer to a subject having a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), a localized or metastatic cancer, or a non-malignant disease. A "true positive" can refer to a subject having a condition and is identified as having the condition by an assay or method of the present disclosure. As used herein, the term "true negative" (TN) refers to a subject who does not have a pathology or has no detectable pathology. A true negative can refer to a subject who does not have a disease or detectable disease, such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), a localized or metastatic cancer, a non-malignant disease, or an otherwise healthy subject. A true negative can refer to a subject who does not have a pathology, has no detectable pathology, or is identified as not having a pathology by an assay or method of the present disclosure.

[0132] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome, whether partial or complete, of any organism or virus that can be used to reference sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz ("UCSC"). "Genome" refers to the complete genetic information of an organism or virus represented by nucleic acid sequences. As used herein, a reference sequence or genome is often an assembled or partially assembled genomic sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. A reference genome can be considered a representative example of a certain set of genes. In some embodiments, a reference genome consists of sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0133] As used herein, the term "sequence reads" or "reads" refers to nucleotide sequences generated by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment ("single-end reads") or from both ends of a nucleic acid (e.g., paired-end reads, double-end reads). In some embodiments, sequence reads (e.g., single-end or paired-end reads) can be generated from one or both strands of a target nucleic acid fragment. The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are an average, median, or average length of about 15 bp to 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 450 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, the sequence reads are an average, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that are less variable, e.g., most of the sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides).For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides in the entire nucleic acid fragment. Sequence reads can be obtained using a variety of methods, such as sequencing techniques, or probes, such as hybridization arrays or capture probes, or amplification techniques, such as polymerase chain reaction (PCR), or linear amplification using a single primer, or isothermal amplification.

[0134] As used herein, terms such as "sequencing" generally refer to any biochemical process that can be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.

[0135] As used herein, the term "sequencing depth" is used interchangeably with the term "coverage" and refers to the number of times a locus is covered by consensus sequence reads corresponding to unique nucleic acid target molecules aligned to the locus, e.g., sequencing depth is equal to the number of unique nucleic acid target molecules covering the locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as "Yx", e.g., 50x, 100x, etc., where "Y" refers to the number of times a locus is covered by sequences corresponding to a nucleic acid target, e.g., the number of independent sequence information covering a particular locus is obtained. In some embodiments, sequencing depth corresponds to the number of genomes sequenced. Sequencing depth can also be applied to multiple loci, or to the entire genome, in which case Y can refer to the average or mean number of times a locus, haploid genome, or entire genome has been sequenced, respectively. Where average depths are quoted, the actual depths at different loci in the dataset may span a range of values. Ultra-deep sequencing refers to a sequencing depth at a locus of at least 100x.

[0136] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.

[0137] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers that are indicative of cancer.

[0138] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protists. Human or non-human animals can be mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), swine (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female of any stage (e.g., man, woman, or child). The subject from whom a sample is taken or who is treated by any of the methods or compositions described herein can be of any age, and can be an adult, an infant or a child.

[0139] As used herein, the term "tissue" can correspond to a group of cells that are together as a functional unit. There can be multiple types of cells in one tissue. Different types of tissues can be composed of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues of different organisms (maternal vs. fetal), or healthy vs. tumor cells. The term "tissue" can generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the term "tissue" or "tissue type" can be used to refer to the tissue from which the cell-free nucleic acid is derived. In one example, the viral nucleic acid fragments can be derived from blood tissue. In another example, the viral nucleic acid fragments can be derived from tumor tissue.

[0140] As used herein, the term "genomic" refers to a characteristic of the genome of an organism. Examples of genomic characteristics include, but are not limited to, those related to the primary nucleic acid sequence of all or a portion of the genome (e.g., the presence or absence of nucleotide polymorphisms, indels, sequence rearrangements, mutation frequencies, etc.), the copy number of one or more specific nucleotide sequences in the genome (e.g., copy number, allele frequency fraction, ploidy of a single chromosome or the entire genome, etc.), the epigenetic state of all or a portion of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome positioning, etc.), and the expression profile of the genome of an organism (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).

[0141] The terms used in this specification are for the purpose of explaining specific cases and are not intended to be limiting. As used in this specification, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, as long as the terms "including", "includes", "having", "has", "with", or their variations are used in any of the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising".

[0142] <Overview of I.C. Cancer Classification> Generally, in cancer classification, a biological sample is sequenced and analyzed to predict cancer from the sequence reads of the genetic material in the biological sample. The workflow can include the actions of one or more physical entities, such as, for example, medical personnel, a sequencing device, an analysis system, etc. The purpose of the workflow includes the detection and / or monitoring of an individual's cancer. From a healthcare perspective, the workflow can play a role in complementing other existing cancer diagnostic tools. The workflow can play a role in providing early detection and / or routine cancer monitoring of an individual diagnosed with cancer in order to better inform the treatment plan for that individual. A general workflow can alternatively be applied more generally to disease classification.

[0143] A medical practitioner performs the sample collection. An individual presents to a medical institution to undergo cancer classification. A medical practitioner collects a sample for performing cancer classification. Examples of biological samples include, but are not limited to, a subject's tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. The sample contains genetic material belonging to the individual and may be extracted and sequenced for cancer classification. Once the sample is collected, it is provided to a sequencing device. Along with the sample, the medical practitioner may collect other information related to the individual, such as biological sex, age, ethnicity, smoking status, past diagnoses, etc. In one or more embodiments, the medical practitioner may utilize a treatment kit. The treatment kit may include one or more sample collection containers. The treatment kit may further include reagents, probes, computer program products, instructions, etc. for use in processing and analyzing the sample.

[0144] A sequencing device performs sample sequencing on the sample. A clinician in a laboratory may perform one or more processing steps on the sample in preparation for sequencing. The clinician may also utilize a processing kit that includes reagents, probes, etc. Once preparation is complete, the clinician loads the sample onto the sequencing device. Examples of devices used for sequencing are further described in connection with Figures 6A and 6B. A sequencing device generally extracts and isolates fragments of nucleic acid to be sequenced to determine the sequence of nucleic acid bases corresponding to the fragments. Sequencing may also include amplification of the nuclear material. Various sequencing processes include Sanger sequencing, fragment analysis, and next generation sequencing. Sequencing may be whole genome sequencing or targeted sequencing using a targeted panel. In the context of DNA methylation, bisulfite sequencing (e.g., as further described in Figures 3A and 3B) can determine the methylation status through bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing provides sequences of multiple nucleic acid fragments in a sample. In one or more embodiments, the sequence may include methylation status vectors, each methylation status vector describing the methylation status for a CpG site on the fragment.

[0145] The analysis system then processes the sequence reads to generate a cancer prediction.

[0146] The analysis system can perform pre-analysis processing, including but not limited to de-duplication of sequence reads, determining metrics for coverage, determining whether a sample is contaminated, removing contaminating fragments, calling sequencing errors, etc. Samples determined to be contaminated may be withheld from further analysis. In other words, the analysis system can withhold further analysis (such as disease or cancer classification) from contaminated samples. In some embodiments, contaminated samples may be physically discarded.

[0147] The analysis system performs one or more analyses. The analysis is a statistical analysis or the application of one or more trained models, and predicts at least the cancer status of the individual from which the sample is derived. Various genetic features can be evaluated and considered, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of genetic mutations. In the context of methylation, the analysis can include identification of abnormal methylation (further described, for example, in FIGS. 4A and 4B), feature extraction (further described, for example, in FIGS. 5A and 5B), and application of a cancer classifier to determine cancer prediction (further described, for example, in FIGS. 5A and 5B). The cancer classifier inputs the features extracted to determine cancer prediction. The cancer prediction is a label or a value. The label may indicate a specific cancer status. For example, a binary label can indicate the presence or absence of cancer, and a multi-class label can indicate one or more cancer types from among multiple cancer types being screened. The value can indicate the likelihood of a specific cancer status, such as the likelihood of cancer, and / or the likelihood of a specific cancer type. The predicted value may further indicate quantification of a cancer signal, which may include quantification of signals from one or more specific origin tissues.

[0148] The analysis system returns the prediction to a healthcare provider. The healthcare provider can establish or adjust a treatment plan based on the cancer prediction. Optimization of treatment is further described in Section VI.C. Treatment.

[0149] [II. Sample Processing] <II.A. Contamination Detection> FIG. 1 is an exemplary flow chart illustrating a process 100 for detecting contamination in a sample, according to one or more embodiments. In general, the sample may be from an individual who is healthy, known to have cancer, suspected to have cancer, or no prior information is known. The sample may be selected from the group consisting of blood, plasma, serum, urine, feces, and saliva samples. Alternatively, the sample may be selected from the group consisting of whole blood, blood fraction (e.g., white blood cells (WBC)), tissue biopsy, pleural fluid, pericardial fluid, cerebrospinal fluid, and peritoneal fluid. Prior to cancer classification or other classification / analysis using the sample, the contamination detection process 100 determines the contamination level of the specimen. For example, the process 100 may output a contamination estimate and / or confidence interval that, when compared to a rule (e.g., a threshold contamination level or interval), determines whether the sample is contaminated. Contamination detection utilizes a set of gene sequences as contamination markers to identify fragments with alleles that are different from the individual's homozygous alleles. Although process 100 is described herein as being performed by an analysis system (examples of which are provided in Figures 6A and 6B and corresponding description), some or all of the steps may be performed by other comparably described sequencing devices and / or computer processors.

[0150] The analysis system sequences the cfDNA fragments in the sample with a target panel that includes a contamination marker probe. The contamination marker is a gene sequence in the human genome and may include an insertion-deletion (indel) site and a plurality of single nucleotide polymorphism (SNP) sites. In some embodiments, the contamination marker is composed of any combination of a plurality of SNP sites listed in Table 1 and an indel site listed in Table 3. For example, the contamination marker includes at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 sites of a plurality of SNP sites listed in Table 1 and at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 sites of an indel site listed in Table 3. The contamination marker probe can be designed to target a contamination marker sequence on a single strand of DNA or a contamination marker sequence on both strands of DNA. In the context of methylation bisulfite sequencing, probes can be designed to target hypermethylated fragments, hypomethylated fragments, or both. Probes can also be designed to target one, some, or all haplotypes of contaminant marker sequences. For example, in the case of multiple SNP sites consisting of two SNP sites (up to four potential haplotypes), probes can target one, some, or all haplotypes. By designing probes that target all of the potential haplotypes of contaminant markers, reference bias to any one particular haplotype can be avoided. In some embodiments, the probes designed for contaminant markers consist of any combination of probes for multiple SNP sites listed in Table 2 and probes for indel sites listed in Table 4. The selection of contaminant markers and subsequent probe design are described below in Figures 2A and 2B. As a result of sequencing, the analysis system has sequence reads of cfDNA fragments in the sample.

[0151] The analysis system identifies 120 or more contaminant markers from the plurality of contaminant markers for which the sample has a homozygous haplotype. To determine the sample's haplotype at each contaminant marker, the analysis system evaluates the haplotypes of all sequence reads of cfDNA fragments at that contaminant marker. The haplotype of the sequence reads can be more accurately determined by realigning the reads to all probes designed for the contaminant marker sites and identifying the alleles corresponding to the probes with the best alignment. If the analysis system observes a sufficient number of reads, e.g., at least 10, 20, 30, 40, 50, and the proportion of sequence reads with one haplotype is very high, e.g., greater than 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, the analysis system determines that the sample is homozygous for that haplotype (i.e., both copies of the contaminant marker in the sample are of the same haplotype). Contaminating markers that do not represent more than a high percentage may be determined to not be homozygous, but rather heterozygous (ie, the two copies of the contaminating marker in the sample are of different haplotypes).

[0152] The analysis system identifies as a contaminating fragment any cfDNA fragment that has a haplotype at one of the identified contaminating markers that is different from the homozygous haplotype of the respective contaminating marker. For each identified contaminating marker, the analysis system labels or otherwise identifies the cfDNA fragment as contaminated (also referred to as "contaminated fragment(s)") if it has a haplotype at the identified contaminating marker that is different from the homozygous haplotype of the sample (step 130). Of the multiple contaminating markers utilized, any sample is likely to have a homozygous haplotype for a subset of the first multiple.

[0153] The analysis system estimates the contamination level based on the identified contaminant fragments (step 140). For example, if no contaminant fragments are identified, the contamination level is 0 or 0%. The analysis system can further estimate the contamination level based on the sequencing depth of the sample, the total number of cfDNA fragments in the sample, the number of contaminant markers performed, the total number of contaminant fragments identified, or a combination thereof. In some examples, the analysis system counts cfDNA fragments with different haplotypes to estimate contamination within a certain confidence interval.

[0154] The analysis system determines whether the sample is contaminated by comparing the contamination level to a threshold (step 150). In some examples, the analysis system compares the estimated contamination level to a threshold contamination level, limit value, or interval. The threshold contamination level or interval can be adjusted or varied depending on the application. For example, a highly sensitive application may require a very low contamination threshold. By way of example only, the threshold contamination level can be a value or interval between 0.1-1.0% or 0.01-1.0%. In some examples, the threshold contamination limit is 0.1%. In some examples, the threshold contamination limit is 0.01%. In some examples, if the estimated contamination level of the sample exceeds the threshold contamination limit value, the analysis system can refuse to generate a cancer classification result (e.g., refuse to call the sample cancer or non-cancer), refuse to include the sample in a training set for training a classifier (e.g., a binary or multi-class classifier for calling cancer / non-cancer or various tissues of origin of cancer), and / or prevent a disease or non-disease classification result based on a contaminated sample. Additionally and / or alternatively, the analytical system may use the comparison to determine whether the sample is contaminated. For example, if the sample is below a threshold, the analytical system may determine that the sample is not contaminated. Conversely, if the sample is above the threshold, the analytical system may determine that the sample is contaminated.

[0155] 2A is an exemplary flow chart illustrating a process 200 for identifying multiple SNP sites for use as contamination markers in contamination detection, according to one or more embodiments. Process 200 is described herein as being performed by an analysis system, although some or all of the steps may be performed by other comparably described sequencing devices and / or computer processors.

[0156] The analysis system identifies multiple SNP sites that are within a threshold distance (step 205). The multiple SNP sites are gene sequences that include at least two SNP sites. In one or more embodiments, the multiple SNP sites include 2, 3, 4, or 5 SNP sites. The analysis system sets a threshold distance, for example, 5 base pairs (bp), 10 bp, 15 bp, 20 bp, or 25 bp. The closer the SNP sites are, the higher the probability that they are present on a single fragment, but a smaller threshold distance also limits the number of viable multiple SNP sites. The analysis system can adjust the threshold distance taking into account the above considerations and / or the budget for contaminant marker probes that can be included in the target assay panel.

[0157] The analysis system includes a plurality of SNP sites having haplotypes with population haplotype frequencies within a threshold range (step 210). At a 2-SNP site, there are two SNPs, each with two variant alleles: a first haplotype at [0, 0] where both SNP sites have no substitutions, a second haplotype at [1, 1] where both SNP sites have substitutions, a third haplotype at [0, 1] where only the second SNP site has a substitution, and a fourth haplotype at [1, 0] where only the first SNP site has a substitution. The analysis system identifies, from the plurality of SNP sites included in step 210, those having two haplotypes (out of four potential haplotypes at the 2-SNP site) whose population haplotype frequencies are within a threshold range of about 50%. The population haplotype frequencies can be obtained from a genomic database or determined using a sample set representative of the population. The threshold range can be ±1%, ±2%, ±3%, ±4%, ±5%, ±6%, ±7%, ±8%, ±9%, or ±10% of 50%. For example, the analysis system includes a 2-SNP site where haplotypes [0,0] and [1,1] have substantial population haplotype frequencies within the threshold range, while [1,0] and [0,1] have small population haplotype frequencies. In another example, the analysis system includes a 2-SNP site where haplotypes [1,0] and [0,1] have substantial population haplotype frequencies within the threshold range, while [0,0] and [1,1] have small population haplotype frequencies.

[0158] The analysis system excludes SNP sites that have guanine-adenine and cytosine-thymine polymorphisms (step 215). The analysis system excludes such SNP sites to avoid problems with bisulfite sequencing, for example, for methylation sequencing. In embodiments that rely on other types of sequencing, the analysis system can skip step 215.

[0159] The analysis system ensures that the haplotypes at each of the plurality of SNP sites are in Hardy-Weinberg equilibrium (step 220). For each SNP site of the plurality of SNP sites, the analysis system calculates the allele frequency of each allele at the SNP site. The allele frequencies can be calculated using data obtained from a genetic database or a sample set representative of a population. The analysis system then checks whether the allele frequencies at each SNP site are in Hardy-Weinberg equilibrium.

number

[0160] The analysis system may omit one or more of steps 205, 210, 215, and 220. As mentioned above, in embodiments that do not perform bisulfite sequencing, step 215 may be omitted. In other embodiments, step 210 and / or step 220 may be omitted.

[0161] At this point, the analysis system has identified SNP sites as possible contaminant markers after one, some, or all of steps 205, 210, 215, and 220. The analysis system can further prune the number of feasible SNP sites based on the budget of contaminant markers that can be implemented in the sequencing panel. The analysis system can optimize the distribution of the multi-SNP site contaminant markers across the genome. The analysis system can also adjust various parameters in steps 205, 210, 215, and 220 to optimize the number of SNP sites selected as contaminant markers. For example, the analysis system can increase the threshold distance in step 205 to increase the number of SNP sites that can be considered in steps 210, 215, and 220. As another example, the analysis system can decrease the threshold range in step 210 to decrease the number of SNP sites that can be considered in steps 215 and 220. Table 1 includes a list of SNP sites selected for use as contaminant markers according to an exemplary embodiment. In one or more embodiments, the SNP sites can be ranked according to a list of criteria. Criteria may include (1) the sequence complexity surrounding the SNP location (k-mer entropy), (2) the similarity of the designed probe to other regions in the genome, (3) the deviation of the population haplotype frequency from an ideal value of 0.5, and (4) the rate of read duplication at that site observed in the actual sequenced samples.

[0162] The analysis system designs a contaminant marker probe targeting each haplotype of the multiple SNP site contaminant markers (step 225). For each multiple SNP site contaminant marker, depending on which two haplotypes are considered in step 210, the analysis system designs a probe targeting each of the two haplotypes. The analysis system can also design a probe targeting both DNA strands of each haplotype. By designing a probe targeting each haplotype of the contaminant marker, reference or substitution bias in sequencing can be avoided. In another embodiment, the analysis system designs a single probe targeting the reference sequence of each multiple SNP site contaminant marker. Table 2 includes a list of probes designed for the multiple SNP site contaminant markers of Table 1 according to an exemplary embodiment.

[0163] 2B is an exemplary flow chart illustrating step 230 of identifying indel sites for use as contamination markers in contamination detection, according to one or more embodiments. Although step 230 is described herein as being performed by an analysis system, some or all of the steps may be performed by other comparably described sequencing devices and / or computer processors.

[0164] The analysis system identifies indel sites within the length range (step 235). An indel site is a genetic sequence that is inserted or deleted from the genome of the individual. The length range can be, for example, 5-10 bp, 5-15 bp, 5-20 bp, 5-25 bp, 5-50 bp, 5-100 bp, 10-15 bp, 10-20 bp, 10-25 bp, 10-50 bp, 10-100 bp, 15-20 bp, 15-25 bp, 15-50 bp, or 15-100 bp.

[0165] The analysis system includes indel sites with high complexity (step 240). Indel sites with low complexity include homopolymers and simple tandem repeats. Homopolymers are strings of one repeated nucleotide, for example, ACGTTTTTTTTTTTTTACG contains a homopolymer of 15 thymines. For homopolymers, there may be a threshold number of repeats to be considered low complexity, for example, 5 or more repeats are considered low complexity. Simple tandem repeats are strings of repeated nucleotide tandems, for example, ACGTCATCATCATCATCATCATACGT contains 7 repeated instances of the nucleotide tandem CAT. For simple tandem repeats, there may be a threshold number of repeats to be considered low complexity, for example, 3 or more repeats are considered low complexity. High complexity sequences include longer sequences that do not contain homopolymers or simple tandem repeats. For example, a high complexity sequence context may include a particularly long indel or a particularly specific long insertion, such as ACGTACCGGGTTTTCA (where "ACCGGGTTTT" is the insertion sequence). The sequence of the high complexity indel ensures screening against contaminating fragments, as opposed to errors introduced by sample processing, such as polymerase chain reaction (PCR) or sequencing.

[0166] The analysis system includes indel sites having population allele frequencies within a threshold range (step 245). The population allele frequencies can be obtained from a genome database or determined using a sample set representative of the population. The threshold range can be ±1%, ±2%, ±3%, ±4%, ±5%, ±6%, ±7%, ±8%, ±9%, or ±10% of 50%. In the case of an insertion sequence, the first allele does not contain the insertion sequence, while the second allele contains the insertion sequence. In the case of a deletion sequence, the first allele contains the deletion sequence, while the second allele does not contain the deletion sequence.

[0167] The analysis system ensures that the alleles at each indel site are in Hardy-Weinberg equilibrium (step 250). For each indel site, the analysis system calculates the allele frequency of each allele at the indel site. The allele frequencies can be calculated using data obtained from a genetic database or a sample set representative of a population. The analysis system then checks whether the allele frequencies at each indel site are in Hardy-Weinberg equilibrium.

number

[0168] The analysis system may omit one or more of steps 235 , 240 , 245 , and 250 .

[0169] At this point, the analysis system has identified indel sites that can be used as contaminant markers after one, some, or all of steps 235, 240, 245, and 250. The analysis system can further reduce the number of viable indel sites based on the budget of contaminant markers that can be implemented in the sequencing panel. The analysis system can optimize the distribution of indel site contaminant markers across the genome. The analysis system can also adjust various parameters in steps 235, 240, 245, and 250 to optimize the number of indel sites selected as contaminant markers. For example, the analysis system can expand the length range in step 235 to increase the number of possible indel sites considered in steps 240, 245, 250. Table 3 includes a list of indel sites selected for use as contaminant markers according to an exemplary embodiment. In one or more embodiments, the indel sites can be ranked according to one or more criteria. Criteria include (1) similarity of the designed probe to other regions of the genome, (2) deviation of the population haplotype frequency from an ideal value of 0.5, (3) read duplication rate at the site observed in the actual sequenced samples, other criteria, or a combination of them. Indel sites can be selected according to ranking and according to a budget for indel sites.

[0170] The analysis system designs 255 contaminant marker probes that target each allele of the indel site contaminant marker. The analysis system designs a probe that targets each of the two alleles. The analysis system may also design a probe that targets both DNA strands of each allele. By designing a probe that targets each allele of the contaminant marker, reference or substitution bias in sequencing can be avoided. In another embodiment, the analysis system designs a single probe that targets the reference sequence of each indel site contaminant marker. Table 4 includes a list of probes designed for the indel site contaminant markers of Table 3 according to an exemplary embodiment.

[0171] <II.B. Generation of Methylation State Vector of DNA Fragments> Figure 3A is an exemplary flowchart illustrating process 300 for sequencing fragments of cfDNA to obtain a methylation state vector, according to one or more embodiments. To analyze DNA methylation, the analysis system first obtains a sample from an individual consisting of a plurality of cfDNA molecules (step 310). In further embodiments, process 300 may also be applied to the sequencing of other types of DNA molecules.

[0172] From the sample, the analysis system can isolate each cfDNA molecule. The cfDNA molecules can be processed to convert unmethylated cytosine to uracil. In one embodiment, this method uses bisulfite treatment of DNA that converts unmethylated cytosine to uracil without converting methylated cytosine. For example, commercially available kits such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct or EZ DNA Methylation™-Lightning kits (available from Zymo Research Corp, Irvine, California) are used for bisulfite conversion. In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, commercially available kits for the conversion of unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, Massachusetts), can be used.

[0173] A sequencing library can be prepared from the converted cfDNA molecules (step 330). During library preparation, unique molecular identifiers (UMIs) can be added to nucleic acid molecules (e.g., DNA molecules) by adapter ligation. UMIs can be short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments (e.g., DNA molecules fragmented by physical shearing, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. UMIs can be degenerate base pairs that function as unique tags that can be used to identify sequence reads derived from specific DNA fragments. In PCR amplification after adapter ligation, the UMIs are replicated along with the bound DNA fragments. This allows for identification of sequence reads derived from the same original fragment in downstream analyses.

[0174] Optionally, the sequencing library can be enriched for cfDNA molecules or genomic regions that are informative for cancer status using multiple hybridization probes. Hybridization probes are short oligonucleotides that can hybridize to specifically identified cfDNA molecules or target regions, enriching those fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to perform deep analysis targeting a specific set of CpG sites of interest to the researcher. Hybridization probes can be 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or more than 10x coverage and tiled across one or more target sequences. For example, hybridization probes tiled with 2x coverage consist of overlapping probes such that each portion of the target sequence hybridizes to two independent probes. Hybridization probes can be tiled across one or more target sequences with less than 1x coverage.

[0175] In one embodiment, hybridization probes are designed to enrich for DNA molecules that have been treated (e.g., using bisulfite) to convert unmethylated cytosines to uracils. During enrichment, hybridization probes (also referred to herein as "probes") can be used to target and pull down nucleic acid fragments that are informative about the presence or absence of cancer (or disease), the state of the cancer, or the classification of the cancer (e.g., cancer class or tissue of origin). Probes can be designed to anneal (or hybridize) to a target strand (complementary) of DNA. The target strand can be the "positive" strand (e.g., the strand that is transcribed into mRNA and then translated into protein) or the complementary "negative" strand. Probes can be 10, 100, or 1000 base pairs in length. Probes can be designed based on a methylation site panel. Probes can be designed based on targeted gene panels to analyze target regions of the genome (e.g., in humans or other organisms) suspected to correspond to specific mutations or specific cancer or other types of disease. Additionally, probes can cover overlapping portions of the target regions.

[0176] Once prepared, the sequencing library or a portion thereof may be sequenced to obtain a plurality of sequence reads. The sequence reads may be in a computer-readable digital format for processing and interpretation by computer software. The sequence reads may be aligned to the reference genome to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information may also include a sequence read length, which may be determined from the start and end positions. The region in the reference genome may relate to a gene or a segment of a gene. The sequence reads are composed of read pairs, denoted as R1 and R2. For example, the first read R1 is sequenced from a first end of a nucleic acid fragment, and the second read R2 is sequenced from a second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 may be consistently aligned (e.g., in opposite orientation) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R1 and R2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R1) and an end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome may represent the likely positions in the reference genome to which the nucleic acid fragment corresponds. An output file in SAM (sequence alignment map) format or BAM (binary) format is generated and output for further analysis, such as determining the methylation status.

[0177] From the sequence reads, the analysis system determines the location and methylation state of each CpG site based on the alignment with the reference genome (step 350). The analysis system generates a methylation state vector for each fragment (step 360) that specifies the location of the fragment in the reference genome (e.g., as specified by the location of the first CpG site of each fragment, or other similar indicator), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, either methylated (e.g., designated as M), unmethylated (e.g., designated as U), or indeterminate (e.g., designated as I). Observed states include methylated and unmethylated states, while unobserved states are indeterminate. Indeterminate methylation states may result from base sequence errors and / or mismatches in the methylation state of the complementary strand of the DNA fragment. The methylation state vectors can be stored in temporary or permanent computer memory for later use and processing. Additionally, the analysis system can remove duplicate reads or duplicate methylation state vectors from a single sample. The analysis system may determine that particular fragments having one or more CpG sites have an uncertain methylation state above a threshold number or percentage, and may exclude such fragments or may selectively include such fragments, but may construct a model that takes such uncertain methylation states into account, one such model being described below in connection with Figure 4.

[0178] 3B is an exemplary diagram of the process 300 of FIG. 3A for sequencing a cfDNA molecule to obtain a methylation status vector, according to one or more embodiments. As an example, the analysis system receives a cfDNA molecule 312, which in this example includes three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 312 are methylated 314. During processing step 320, the cfDNA molecule 312 is converted to generate a converted cfDNA molecule 322. During processing step 320, the second CpG site, which was not methylated, has its cytosine converted to uracil. However, the first and third CpG sites were not converted.

[0179] After conversion, a sequencing library is prepared (step 330) and sequencing is performed (step 340) to generate sequence reads 342. The analysis system aligns the sequence reads 342 to a reference genome 344 (step 350). The reference genome 344 provides context for where in the human genome the fragment cfDNA originates. In this simplified example, the analysis system aligns the sequence reads 342 such that three CpG sites correlate to CpG sites 23, 24 and 25 (arbitrary reference identifiers used for convenience of illustration) (step 350). In this way, the analysis system can generate information on both the methylation status of all CpG sites on the cfDNA molecule 312 and the location in the human genome to which the CpG sites map. As shown, CpG sites on the sequence reads 342 that are methylated are read as cytosines. In this example, cytosine appears only at the first and third CpG sites in sequence read 342, so it can be inferred that the first and third CpG sites in the original cfDNA molecule are methylated. However, the second CpG site can be read as a thymine (U is converted to T during sequencing), so it can be inferred that the second CpG site is unmethylated in the original cfDNA molecule. Using these two pieces of information, methylation status and position, the analysis system generates a methylation status vector 352 for fragment cfDNA 312 (step 360). In this example, the resulting methylation status vector 352 is <M 23 , U 24 , M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the position of each CpG site in the reference genome.

[0180] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in biological samples. The one or more sequencing methods can include any form of sequencing that can be used to obtain a measured number of sequence reads from nucleic acids (e.g., cell-free nucleic acids), including the Roche 454 platform, Applied Biosystems SOLID platform, Helicos True Single Molecule DNA sequencing technology, Affymetrix's sequencing-by-hybridization platform, Pacific Biosciences' single molecule real-time (SMRT) technology, 454 Life Sciences, Illumina / Solexa, Helicos Biosciences' sequencing-by-synthesis platform, and Applied Biosystems' sequencing-by-ligation platform. Life technologies' ION TORRENT technology and Nanopore sequencing can also be used to obtain sequence reads from nucleic acids (e.g., cell-free nucleic acids) in biological samples. Using sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer, Genome Analyzer II, HISEQ 2000, HISEQ 4500 (Illumina, San Diego, CA)), sequence reads can be obtained from cell-free nucleic acids obtained from biological samples of training subjects to form a genotype dataset. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. One example of this type of sequencing technology uses a flow cell that contains an optically clear slide, with eight individual lanes on its surface to which oligonucleotide anchors (e.g., adapter primers) are attached. Cell-free nucleic acid samples can contain signals or tags to facilitate detection.Obtaining sequence reads from cell-free nucleic acids obtained from a biological sample can include obtaining quantitative information of signals or tags using a variety of techniques, such as, for example, flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarrays, mass spectrometry, cytofluorimetry analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and combinations thereof.

[0181] The one or more sequencing methods may include a whole genome sequencing assay. The whole genome sequencing assay may include a physical assay that generates sequence reads for the whole genome or a substantial part of the whole genome, which may be used to determine large variations such as copy number variations or copy number abnormalities. Such physical assays may use whole genome sequencing technology or whole exome sequencing technology. The whole genome sequencing assay may have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the whole genome of the subject. In some embodiments, the sequencing depth is about 30,000x. The one or more sequencing methods may include a targeted panel sequencing assay. The targeted panel sequencing assay may have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x for a targeted panel of genes. A gene targeting panel may include between 450 and 500 genes. A gene targeting panel may include a range of 500±5 genes, a range of 500±10 genes, or a range of 500±25 genes.

[0182] The one or more sequencing methods may include paired-end sequencing. The one or more sequencing methods may generate multiple sequence reads. The multiple sequence reads may have an average length ranging from 10 to 700, 50 to 400, or 100 to 300. The one or more sequencing methods may include a methylation sequencing assay. The methylation sequencing may be i) whole genome methylation sequencing, or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, the methylation sequencing is whole genome bisulfite sequencing (e.g., WGBS). The methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes targeting the most informative regions of the methylome, a proprietary methylation database, and a prior prototypic whole genome and targeted sequencing assay.

[0183] Methylation sequencing may detect one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each nucleic acid methylation fragment. Methylation sequencing may include converting one or more unmethylated cytosines or one or more methylated cytosines to one or more corresponding uracils in each nucleic acid methylation fragment. One or more uracils may be detected as one or more corresponding thymines during methylation sequencing. The conversion of one or more unmethylated cytosines or one or more methylated cytosines may consist of chemical conversion, enzymatic conversion, or a combination thereof.

[0184] For example, bisulfite conversion involves converting cytosines to uracils while leaving methylated cytosines (e.g., 5-methylcytosine or 5-mC) intact. In some DNA, about 95% of cytosines may be unmethylated, and the resulting DNA fragments may contain many uracils represented by thymines. Enzymatic conversion processes can be used to treat nucleic acids prior to sequencing, which can be done in a variety of ways. One example of bisulfite-free conversion consists of TET-assisted pyridine borane sequencing (TAPS), a bisulfite-free base resolution sequencing method for non-destructive and direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosines. The methylation status of the CpG sites in the corresponding multiple CpG sites in each nucleic acid methylation fragment may be methylated if the CpG site is determined to be methylated by methylation sequencing, or unmethylated if the CpG site is determined to be unmethylated by methylation sequencing.

[0185] Methylation sequencing assays (e.g., WGBS and / or targeted methylation sequencing) may have an average sequencing depth, including but not limited to about 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or up to 30,000x. Methylation sequencing may have a sequencing depth of more than 30,000x, for example, at least 40,000x or 50,000x. Whole genome bisulfite sequencing may have an average sequencing depth of 20x to 50x, and targeted methylation sequencing may have an average effective depth of 100x to 1000x, where the effective depth may be the equivalent whole genome bisulfite sequencing coverage to obtain the same number of sequence reads obtained by targeted methylation sequencing.

[0186] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, for example, U.S. Patent Application No. 62 / 642,480, entitled "Methylation Fragment Anomaly Detection," filed March 13, 2018, and U.S. Patent Application No. 16 / 719,902, entitled "Systems and Methods for Estimating Cell Source Fractions Using Methylation Information," filed December 18, 2019, each of which is incorporated herein by reference. Other methods for methylation sequencing, including those disclosed herein and / or any modifications, permutations, or combinations thereof, can be used to obtain fragment methylation patterns. Methylation sequencing can be used to identify one or more methylation status vectors, for example, as described in U.S. patent application Ser. No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019, or in accordance with any of the techniques disclosed in U.S. patent application Ser. No. 15 / 931,022, entitled "Model-Based Featurization and Classification," filed May 13, 2020, each of which is incorporated herein by reference.

[0187] Multiple nucleic acid methylation fragments can be obtained using nucleic acid methylation sequencing and one or more methylation state vectors obtained. Each corresponding plurality of nucleic acid methylation fragments (e.g., for each genotype dataset) can be composed of 100 or more nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can consist of 1000 or more nucleic acid methylation fragments, 5000 or more nucleic acid methylation fragments, 10,000 or more nucleic acid methylation fragments, 20,000 or more nucleic acid methylation fragments, or 30,000 or more nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can be 10,000 to 50,000 nucleic acid methylation fragments. The corresponding plurality of nucleic acid methylation fragments can consist of 1000 or more, 10,000 or more, 100,000 or more, 1,000,000 or more, 10,000,000 or more, 100,000,000 or more, 500,000,000 or more, 1,000,000,000 or more, 2,000,000,000 or more, 3,000,000,000 or more, 4,000,000,000 or more, 5,000,000,000 or more, 6,000,000,000 or more, 7,000,000,000 or more, 8,000,000,000 or more, 9,000,000,000 or more, or 10,000,000,000 or more nucleic acid methylation fragments. The average length of the corresponding plurality of nucleic acid methylation fragments can be 140 to 480 nucleotides.

[0188] Methods for sequencing nucleic acids and further details regarding methylation sequencing data are disclosed in U.S. Provisional Patent Application No. 62 / 985,258, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed on March 4, 2020, which is hereby incorporated by reference in its entirety.

[0189] <II.C. Identification of Abnormal Fragments> The analysis system can determine abnormal fragments of a sample using the methylation state vector of the sample. For each fragment in the sample, the analysis system can determine whether the fragment is an abnormal fragment using the methylation state vector corresponding to that fragment. In some embodiments, the analysis system calculates a p-value score that describes, for each methylation state vector, the probability of observing that methylation state vector or another methylation state vector with an even lower probability in a healthy control group. The process of calculating the p-value score will be further described later in the P-value filtering of Section II.C.i. The analysis system can determine a fragment having a methylation state vector below a threshold p-value score as an abnormal fragment. In some embodiments, the analysis system further labels fragments having at least some CpG sites with a methylation or unmethylation ratio exceeding a threshold as hypermethylated and hypomethylated fragments, respectively. Hypermethylated or hypomethylated fragments may sometimes also be referred to as abnormal fragments with extreme methylation (UFXM). In other embodiments, the analysis system can implement various other probability models for determining abnormal fragments. Examples of other probability models include mixture models, deep probability models, and the like. In some embodiments, the analysis system can use any combination of the processes described below to identify abnormal fragments. Using the identified abnormal fragments, the analysis system can filter a set of methylation state vectors of the sample for use in other processes, such as training and deployment of a cancer classifier.

[0190] <II.C.I. P-value Filtering> In some embodiments, the analysis system calculates a p-value score for each methylation status vector compared to the methylation status vector from the fragments of the healthy control group. The p-value score can represent the probability of observing a methylation status that matches the methylation status vector, or the probability of observing other, less likely methylation status vectors in the healthy control group. To determine whether a DNA fragment is abnormally methylated, the analysis system can use a healthy control group in which normally methylated fragments predominate. When performing this probabilistic analysis to determine abnormal fragments, the determination can be weighted in comparison to the control subjects that constitute the healthy control group. To ensure robustness in the healthy control group, the analysis system can select a certain threshold number of healthy individuals to provide samples containing the DNA fragments. Figure 4A below describes how to generate a data structure for a healthy control group from which the analysis system can calculate a p-value score. Figure 4B describes how to calculate a p-value score by using the generated data structure.

[0191] 4A is a flow chart illustrating a process 400 for generating a healthy control data structure, according to an embodiment. To create the healthy control data structure, an analysis system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. A methylation state vector can be identified for each fragment, e.g., via process 300.

[0192] Using the methylation state vector of each fragment, the analysis system can subdivide the methylation state vector into strings of CpG sites. In some embodiments, the analysis system subdivides the methylation state vector (step 405) such that all the resulting strings are less than a predetermined length. For example, subdividing a methylation state vector of length 11 into strings of length 3 or less results in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, subdividing a methylation state vector of length 7 into strings of length 4 or less results in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector is less than or equal to the length of the specified string, the methylation state vector may be converted into a single string that contains all the CpG sites in the vector.

[0193] The analysis system tallies strings (step 410) by counting the number of strings that have the specified CpG site as the first CpG site in the string and have that methylation state possibility in the control group that have that methylation state possibility. For example, at a CpG site, if the string is of length 3, there are 2^3, or 8 possible string configurations. For each of the 8 possible string configurations at that given CpG site, the analysis system tallies how many vector possibilities of each methylation state appear in the control group. Continuing with this example, this means: <M x , M x+1 , M x+2 >, <M x , M x+1 , U x+2 >,..., x , U x+1 , U x+2 The analysis system may include tabulating the calculated counts for each start CpG site x in the reference genome. The analysis system creates a data structure that stores the tabulated counts for each start CpG site and string possibility (step 415).

[0194] ​Setting an upper limit on string length has several advantages. First, depending on the maximum string length, it can dramatically increase the size of the data structures created by the analysis system. For example, a maximum string length of 4 means that every CpG site has at least 2^4 numbers to count for strings of length 4. Increasing the maximum string length to 5 means that every CpG site has an additional 2^4 or 16 numbers to count, doubling the number to count (and the computer memory required) compared to the previous string length. By keeping the string size small, the creation and performance of the data structure (e.g., used to access it later, as described below) can be kept reasonable in terms of computation and storage. Second, as a statistical consideration limiting the maximum string length, overfitting of downstream models that use string counts can be avoided. In cases where long strings of CpG sites do not biologically have a strong impact on the outcome (e.g., predicting anomalies that predict the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic as it uses a significant amount of data that may not be available and therefore may be too sparse for the model to perform adequately. For example, if one were to calculate the probability of abnormality / cancer conditional on a priori 100 CpG sites, one could use counts of strings in a data structure of length 100, ideally with exact matches to the methylation states of the priori 100. If the length 100 strings are only sparsely counted, there may not be enough data to determine whether a length 100 string in a test sample is abnormal or not.

[0195] 4B is a flow chart illustrating step 420 for identifying aberrantly methylated fragments from a sample, according to one or more embodiments. In step 420, the analysis system generates methylation status vectors from the cfDNA fragments of the subject (step 300). The analysis system can process each methylation status vector as follows:

[0196] For a given methylation state vector, the analysis system enumerates all possible methylation state vectors that have the same starting CpG site and the same length (i.e., set of CpG sites) in the methylation state vector (step 430). Because each methylation state is generally either methylated or unmethylated, there are effectively two possibilities for each CpG site, and thus the number of different possibilities for a methylation state vector depends on a power of two, and a methylation state vector of length n has 2 n If the methylation state vector includes an uncertain state for one or more CpG sites, the analysis system can enumerate the possible methylation state vectors by considering only CpG sites with an observed state (step 430).

[0197] The analysis system calculates the probability of observing each possibility of the methylation state vector by accessing the healthy control group data structure for the identified starting CpG site and methylation state vector length (step 440). In some embodiments, calculating the probability of observing a given possibility uses Markov chain probabilities to model the joint probability calculation. The Markov model can be trained based at least in part on the evaluation of the methylation state of each CpG site at the corresponding multiple CpG sites of each fragment (e.g., nucleic acid methylation fragment) across those nucleic acid methylation fragments in the healthy non-cancer cohort dataset with the corresponding multiple CpG sites. For example, a Markov model (e.g., a hidden Markov model or HMM) is used to determine the probability that a sequence of methylation states (e.g., consisting of "M" or "U") may be observed for a nucleic acid methylation fragment in the multiple nucleic acid methylation fragments, given a set of probabilities that determine the likelihood of observing the next state in the sequence for each state in the sequence. The set of probabilities can be obtained by training an HMM. Such training may include calculating statistical parameters (e.g., the probability that a first state may transition to a second state (transition probability) and / or the probability that a given methylation state may be observed for each CpG site (release probability)) given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). HMMs may be trained using supervised training (e.g., using samples where the underlying sequences as well as the observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation maximization training, and / or Baum-Welch training). In other embodiments, computational methods other than Markov chain probabilities are used to determine the probability of observing each possibility of the methylation state vector. For example, such computational methods may include learned representations. The p-value threshold may be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.

[0198] The analysis system uses the calculated probability for each possibility to calculate a p-value score for the methylation status vector (step 450). In some embodiments, this involves identifying a calculated probability that corresponds to the possibility of matching the methylation status vector in question. Specifically, this may be the possibility of having the same set of CpG sites as the methylation status vector, or similarly the same start CpG site and length. The analysis system may sum the calculated probabilities of all possibilities that have a probability less than or equal to the identified probability to generate a p-value score.

[0199] This p-value can represent the probability of observing the fragment methylation state vector, or other methylation state vectors with even lower probability in healthy control group. A low p-value score may generally correspond to a methylation state vector that is rare in healthy individuals relative to the healthy control group, and the fragment is labeled as abnormally methylated. A high p-value score may generally relate to a methylation state vector that is expected to be present in a relative sense in healthy individuals. For example, if the healthy control group is a non-cancer group, a low p-value may indicate that the fragment is abnormally methylated compared to the non-cancer group, and thus indicate the presence of cancer in the subject.

[0200] As described above, the analysis system can calculate a p-value score for each of a plurality of methylation status vectors, each representing a cfDNA fragment in a test sample. To identify which fragments are abnormally methylated, the analysis system can filter the set of methylation status vectors based on the p-value score (step 460). In some embodiments, the filtering is performed by comparing the p-value score to a threshold and retaining only fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, etc.

[0201] According to example results of process 400, the analysis system can obtain a median (range) of 2,800 (1,500-12,000) fragments with aberrant methylation patterns for training participants without cancer and a median (range) of 3,000 (1,200-420,000) fragments with aberrant methylation patterns for training participants with cancer. These filter sets of fragments with aberrant methylation patterns can be used for downstream analysis, as described below in Section III.

[0202] In some embodiments, the analysis system uses a sliding window to determine the probabilities of the methylation state vector and to calculate p-values ​​(step 455). Rather than enumerating the possibilities and calculating a p-value for the entire methylation state vector, the analysis system can enumerate the possibilities and calculate a p-value only for a window of contiguous CpG sites, where the window is shorter in length (of CpG sites) than at least some fragment (otherwise the window serves no purpose). The length of the window can be selected statically, user-determined, dynamically, or in other ways.

[0203] When calculating p-values ​​for a methylation status vector larger than a window, the window may identify a contiguous set of CpG sites from the vector within the window starting from the first CpG site in the vector. The analysis system may calculate a p-value score for the window that includes the first CpG site. The analysis system may then "slide" the window to a second CpG site in the vector and calculate another p-value score for the second window. Thus, for a window size l and a methylation vector length m, each methylation status vector may generate m-l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all the sliding windows may be taken as the overall p-value score for the methylation status vector. In other embodiments, the analysis system may aggregate the p-value scores of the methylation status vectors to generate an overall p-value score.

[0204] Using a sliding window allows the number of enumerated possibilities of methylation state vectors and the corresponding probability calculations to be reduced. A realistic example is a fragment with more than 54 CpG sites. Instead of calculating probabilities for 2^54 (~1.8x10^16) possibilities to generate a single p-score, the analysis system can use a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of the methylation state vector of that fragment. Each of the 50 calculations can enumerate 2^5 (32) possibilities of the methylation state vector, resulting in a total of 50x2^5 (1.6x10^3) ​​probability calculations. The result is a significant reduction in the calculations to be performed, which is not meaningful for the precise identification of abnormal fragments.

[0205] In embodiments with an uncertain state, the analysis system may calculate a p-value score summing the CpG sites with an uncertain state in the methylation state vector of the fragment. The analysis system may identify all possibilities that have a consensus with all methylation states of the methylation state vector excluding the uncertain state. The analysis system may assign a probability to the methylation state vector as the sum of the probabilities of the identified possibilities. As an example, the analysis system may calculate a p-value score summing the CpG sites with an uncertain state in the methylation state vector of the fragment ...<M1、M2、U3> and<M1、U2、U3> As the sum of the probabilities of the possible methylation state vectors of<M1、I2、U3> The probability of a methylation state vector having one or more uncertain states can be calculated. This method of summing CpG sites with uncertain states can use a calculation of the probability of up to 2^i possibilities, where i indicates the number of uncertain states in the methylation state vector. In additional embodiments, a dynamic programming algorithm can be implemented to calculate the probability of a methylation state vector having one or more uncertain states.

[0206] In some embodiments, the computational burden of calculating the probability and / or p-value scores may be further reduced by caching at least some of the calculations. For example, the analysis system may cache the calculation of the probability for the methylation status vector (or a window thereof) in temporary or persistent memory. By caching the probability probabilities, p-score values ​​may be efficiently calculated without the need to recalculate the underlying probability probabilities if other fragments have the same CpG sites. Similarly, the analysis system may calculate a p-value score for each of the possibilities of the methylation status vector associated with a set of CpG sites from the vector (or a window thereof). The analysis system may cache the p-value scores for use in determining the p-value scores of other fragments containing the same CpG sites. In general, the p-value scores of the possibilities of methylation status vectors with the same CpG sites may be used to determine the p-value scores of different ones of the possibilities from the same set of CpG sites.

[0207] The one or more nucleic acid methylation fragments can be filtered before training the region model or the cancer classifier. Filtering the nucleic acid methylation fragments can include removing each nucleic acid methylation fragment that does not meet one or more selection criteria (e.g., below or above one selection criterion) from the corresponding plurality of nucleic acid methylation fragments. The one or more selection criteria can include a p-value threshold. The output p-value of each nucleic acid methylation fragment can be determined at least in part based on a comparison of the corresponding methylation pattern of each nucleic acid methylation fragment with a corresponding distribution of the methylation patterns of the nucleic acid methylation fragments in a healthy non-cancer cohort dataset having the corresponding plurality of CpG sites of each nucleic acid methylation fragment.

[0208] Filtering the plurality of nucleic acid methylation fragments may include removing each nucleic acid methylation fragment that does not satisfy a p-value threshold. A filter may be applied to each methylation pattern of each nucleic acid methylation fragment using the methylation patterns observed across the first plurality of nucleic acid methylation fragments. Each methylation pattern of each nucleic acid methylation fragment (e.g., fragment 1, ..., fragment N) is identified with a methylation site identifier and a corresponding methylation pattern, represented as a sequence of 1's and 0's, and includes one or more corresponding methylation sites (e.g., CpG sites), where each "1" represents a methylated CpG site in the one or more CpG sites and each "0" represents an unmethylated CpG site in the one or more CpG sites. The methylation patterns observed across the first plurality of nucleic acid methylation fragments may be used to construct a methylation state distribution for the CpG site states (e.g., CpG site A, CpG site B, ..., CpG site ZZZ) collectively represented by the first plurality of nucleic acid methylation fragments. Further details regarding the processing of nucleic acid methylation fragments are disclosed in U.S. Provisional Patent Application No. 62 / 985,258, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed March 4, 2020, the entire contents of which are incorporated herein by reference.

[0209] If each nucleic acid methylated fragment has an abnormal methylation score below the abnormal methylation score threshold, each nucleic acid methylated fragment may not satisfy the selection criteria in one or more selection criteria. In this situation, the abnormal methylation score may be determined by a mixture model. For example, the mixture model may detect an abnormal methylation pattern of a nucleic acid methylated fragment by determining the likelihood of a methylation state vector (e.g., a methylation pattern) of each nucleic acid methylated fragment based on the number of possible methylation state vectors of the same length and the same corresponding genomic location. This may be performed by generating multiple possible methylation states for a vector of a specified length at each genomic location of the reference genome. The multiple possible methylation states may be used to determine the total number of possible methylation states and the subsequent probability of each predicted methylation state at the genomic location. The likelihood of a sample nucleic acid methylated fragment corresponding to a genomic location in the reference genome may then be determined by matching the sample nucleic acid methylated fragment to the predicted (e.g., possible) methylation states and retrieving the calculated probability of the predicted methylation state. An abnormal methylation score may then be calculated based on the probability of the sample nucleic acid methylated fragment.

[0210] Each nucleic acid methylated fragment may not satisfy the selection criteria in one or more selection criteria if the number of residues in each nucleic acid methylated fragment is less than a threshold. The threshold number of residues may be 10-50, 50-100, 100-150, or 150 or more. The threshold number of residues may be a fixed value of 20-90. Each nucleic acid methylated fragment may not satisfy the selection criteria in one or more selection criteria if the each nucleic acid methylated fragment has less than a threshold number of CpG sites. The threshold number of CpG sites may be 4, 5, 6, 7, 8, 9, or 10. Each nucleic acid methylated fragment may not satisfy the selection criteria in one or more selection criteria if the genomic start and end positions of each nucleic acid methylated fragment indicate that each nucleic acid methylated fragment represents less than a threshold number of nucleotides in the human genome reference sequence.

[0211] The filtering can remove nucleic acid methylated fragments in the corresponding plurality of nucleic acid methylated fragments that have the same corresponding methylation pattern and the same corresponding genomic start and end positions as another nucleic acid methylated fragment in the corresponding plurality of nucleic acid methylated fragments. This filtering step can remove redundant fragments that are complete duplicates, possibly including PCR duplicates. This filtering can remove nucleic acid methylated fragments that have the same corresponding genomic start and end positions as another nucleic acid methylated fragment in the corresponding plurality of nucleic acid methylated fragments, and the number of different methylation states is less than a threshold. The threshold number of different methylation states used to retain the nucleic acid methylated fragments can be 1, 2, 3, 4, 5, or 5 or more. For example, a first nucleic acid methylated fragment that has the same corresponding genomic start and end positions as a second nucleic acid methylated fragment, but has at least 1, at least 2, at least 3, at least 4, or at least 5 different methylation states at each CpG site (e.g., aligned to a reference genome) is retained. As another example, a first nucleic acid methylated fragment that has the same methylation state vector (eg, methylation pattern) but differs in the genomic start and end locations that correspond to the second nucleic acid methylated fragment is also retained.

[0212] Filtering can remove assay artifacts in multiple nucleic acid methylation fragments. Removal of assay artifacts can include removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences that did not undergo conversion during bisulfite conversion. Filtering can remove contaminants (e.g., resulting from sequencing, nucleic acid isolation, and / or sample preparation).

[0213] Filtering can remove a subset of methylated fragments from a plurality of methylated fragments based on mutual information filtering of each methylated fragment for cancer states across a plurality of training subjects. For example, mutual information can provide a measure of the interdependence between two simultaneously sampled states of interest. Mutual information can be determined by selecting an independent set of CpG sites (e.g., within all or part of a nucleic acid methylated fragment) from one or more data sets and comparing the probabilities of methylation states for the set of CpG sites between two sample groups (e.g., subsets and / or groups of a genotype data set, biological samples, and / or subjects). The mutual information score can indicate the probability of the methylation pattern of a first condition to a second condition in each region of each frame of a sliding window and thus indicate the discriminative power of each region. The mutual information score can be similarly calculated for each region of each frame of the sliding window as it progresses across the selected set of CpG sites and / or the selected genomic region. Further details regarding mutual information filtering are disclosed in U.S. Patent Application No. 17 / 119,606, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed on December 11, 2020, which is hereby incorporated by reference in its entirety.

[0214] <II.C.II. Hypermethylated Fragments and Hypomethylated Fragments> In some embodiments, the analysis system determines an abnormal fragment as a fragment in which the number of CpG sites is equal to or greater than a threshold value, and the methylation rate of the CpG sites is equal to or greater than a threshold value, or the unmethylation rate of the CpG sites is equal to or greater than a threshold value. Examples of the threshold value for the length of the fragment (or CpG site) include 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, and the like. Examples of the threshold value for the percentage of methylation or unmethylation include 80% or more, 85% or more, 90% or more, 95% or more, or other ratios within the range of 50% to 100%.

[0215] <II.D. Exemplary Analysis System> FIG. 6A is an exemplary flowchart of an apparatus for sequencing a nucleic acid sample according to one or more embodiments. This exemplary flowchart includes devices such as a sequencer 620 and an analysis system 600. The sequencer 620 and the analysis system 600 may operate in conjunction to perform one or more steps in the process 300 of FIG. 3A, the process 400 of FIG. 4A, the step 420 of FIG. 4B, and other steps described herein.

[0216] In various embodiments, the sequencer 620 receives the concentrated nucleic acid sample 610. As shown in FIG. 6A, the sequencer 620 may include a graphical user interface 625 that enables user interaction with a specific task (e.g., start of sequencing or end of sequencing), and one or more loading stations 630 for loading a sequencing cartridge containing the concentrated fragment sample and / or for loading the buffers necessary to perform the sequencing assay. Thus, when the user of the sequencer 620 prepares the necessary reagents and the sequencing cartridge at the loading station 630 of the sequencer 620, the user can start the sequencing by interacting with the graphical user interface 625 of the sequencer 620. Once started, the sequencer 620 performs the sequencing and outputs sequence reads of the concentrated fragments from the nucleic acid sample 610.

[0217] In some embodiments, the sequencer 620 is communicatively coupled to the analysis system 600. The analysis system 600 includes several computing devices used to process sequence reads for various applications, such as assessment of methylation status at one or more CpG sites, variant calling, or quality control. The sequencer 620 can provide sequence reads to the analysis system 600 in BAM file format. The analysis system 600 can be communicatively coupled to the sequencer 620 through wireless, wired, or a combination of wireless and wired communication techniques. In general, the analysis system 600 is configured to include a processor and a non-transitory computer-readable storage medium that stores computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein.

[0218] In some embodiments, sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information, for example, via step 340 of process 300 of FIG. 3A. The alignment position may generally describe the start and end positions of a region in the reference genome that corresponds to the starting and ending nucleotide bases of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate the first and last CpG sites included in the sequence read according to the alignment to the reference genome. The alignment position information may further indicate the methylation status and positions of all CpG sites in a given sequence read. Regions in the reference genome may be associated with genes or segments of genes, and as such, the analysis system 600 may label a sequence read with one or more genes that align to the sequence read. In one embodiment, the length (or size) of a fragment is determined from the start and end positions.

[0219] In various embodiments, for example, when a paired-end sequencing process is used, the sequence reads are composed of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be consistently aligned (e.g., in opposite directions) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R_1 and R_2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R_1) and an end position in the reference genome corresponding to the end of the second read (e.g., R_2). In other words, the start position and the end position in the reference genome may represent the likely positions in the reference genome to which the nucleic acid fragment corresponds. Output files in SAM (sequence alignment map) or BAM (binary) format are generated and exported for further analysis.

[0220]

[0036] Referring now to Figure 6B, Figure 6B is a block diagram of an analysis system 600 for processing DNA samples according to one embodiment. The analysis system implements one or more computing devices for use in analyzing DNA samples. Analysis system 600 includes a sequence processor 640, a sequence database 645, a model database 655, a model 650, a parameter database 665, and a score engine 660. In some embodiments, analysis system 600 performs some or all of process 300 of Figure 3A and process 400 of Figure 4A.

[0221] The sequence processor 640 generates methylation state vectors for fragments from the sample. For each CpG site on the fragment, the sequence processor 640 generates a methylation state vector for each fragment, via process 300 of FIG. 3A, that specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, either methylated, unmethylated, or uncertain. The sequence processor 640 can store the fragment methylation state vectors in a sequence database 645. The data in the sequence database 645 may be organized such that the methylation state vectors from the samples are associated with each other.

[0222] Additionally, multiple different models 650 may be stored in a model database 655 or retrieved for use with a test sample. In one embodiment, the model is a cancer classifier trained to determine a cancer prognosis for a test sample using a feature vector derived from the abnormal fragment. Training and use of cancer classifiers is further described in relation to "Section III. Cancer Classifiers for Determining Cancer." Analysis system 600 may train one or more models 650 and store various trained parameters in a parameter database 665. Analysis system 600 stores models 650 along with features in model database 655.

[0223] During inference, the score engine 660 uses one or more models 650 to return outputs. The score engine 660 accesses the models 650 in the model database 655 along with trained parameters from the parameter database 665. According to each model, the score engine receives appropriate inputs for the model and calculates outputs based on the received inputs, parameters, and functions of each model that relate inputs and outputs. In some use cases, the score engine 660 further calculates metrics related to confidence in the outputs calculated from the models. In other use cases, the score engine 660 calculates other intermediate values ​​for use with the models.

[0224] [III. Cancer Classifier for Cancer Detection] <III.A. Overview> The cancer classifier can be trained to receive a feature vector of a test sample and determine whether the test sample is from a subject having cancer, more specifically a particular cancer type. The cancer classifier can include a plurality of classification parameters, a function that operates on the input feature vector using the classification parameters, and a function that represents the relationship between the input feature vector and a cancer prediction as an output determined by the function. In some embodiments, the feature vector input to the cancer classifier is based on a set of abnormal fragments determined from the test sample. The abnormal fragments may be determined via step 420 of FIG. 4B, and more specifically, hypermethylated fragments and hypomethylated fragments determined via step 470 of step 420, or abnormal fragments determined according to some other step. Before deploying the cancer classifier, the analysis system can train the cancer classifier.

[0225] <III.B. Training of Cancer Classifier> FIG. 5A is a flowchart illustrating a process 500 for training a cancer classifier according to an embodiment. The analysis system obtains a plurality of training samples, each having a set of abnormal fragments and a label of a cancer type (step 510). The plurality of training samples can include any combination of samples from healthy subjects having a general label of "non-cancer", samples from subjects having a general label of "cancer", or samples from subjects having a specific label (e.g., "breast cancer", "lung cancer", etc.). Training samples from subjects of a certain cancer type may be referred to as a cohort of that cancer type or a cancer type cohort. The analysis system can ensure that the training samples used for training the cancer classifier are not contaminated. To determine whether the training samples are contaminated, the analysis system can execute the process 100 of FIG. 1.

[0226] The analysis system determines, for each training sample, a feature vector based on the set of abnormal fragments of the training sample (step 520). The analysis system can calculate an anomaly score for each CpG site in the initial set of CpG sites. The initial set of CpG sites can be all CpG sites in the human genome or a portion thereof, and can be up to 10 4 , 10 5 , 10 6 , 10 7 , 10 8 In one embodiment, the analysis system defines an abnormality score for the feature vector with binary scoring based on the presence or absence of an abnormal fragment encompassing a CpG site in the set of abnormal fragments. In another embodiment, the analysis system defines an abnormality score based on a count of abnormal fragments that overlap a CpG site. In one example, the analysis system can use a ternary scoring that assigns a first score if no abnormal fragments are present, a second score if a small number of abnormal fragments are present, and a third score if a small number or more abnormal fragments are present. For example, the analysis system counts five abnormal fragments in a sample that overlap a CpG site and calculates an abnormality score based on the five counts.

[0227] Once all the anomaly scores have been determined for the training samples, the analysis system can determine a feature vector as a vector of elements, each element including one of the anomaly scores associated with one of the CpG sites in the initial set. The analysis system can normalize the anomaly scores of the feature vector based on the coverage of the sample, where coverage can refer to the median or average sequencing depth across all CpG sites covered by the initial set of CpG sites used for the classifier, or based on the set of anomaly fragments for a given training sample.

[0228] As an example, see FIG. 5B, which illustrates a matrix 522 of training feature vectors. In this example, the analysis system has identified a CpG site [K] 526 to consider in generating a feature vector for a cancer classifier. The analysis system selects a training sample [N] 524. The analysis system determines a first abnormality score 528 for a first arbitrary CpG site [k1] used in the feature vector of the training sample [n1]. The analysis system checks each abnormal fragment in the set of abnormal fragments. If the analysis system identifies at least one abnormal fragment that includes the first CpG site, the analysis system determines a first abnormality score 528 for the first CpG site as 1, as shown in FIG. 5B. Considering a second arbitrary CpG site [k2], the analysis system similarly checks a set of at least one abnormal fragment that includes the second CpG site [k2]. If the analysis system finds no such aberration fragment containing the second CpG site, then the analysis system determines a second aberration score 529 for the second CpG site [k2] to be 0, as shown in Figure 5B. Once the analysis system has determined all of the aberration scores for the initial set of CpG sites, the analysis system determines a feature vector for the first training sample [n1] that includes anomaly scores including a first aberration score 528 of 1 for the first CpG site [k1], a second aberration score 529 of 0 for the second CpG site [k2], and so forth, thereby forming a feature vector [1, 0, ...].

[0229] Other approaches to sample characterization are described in U.S. Application No. 15 / 931,022, entitled "Model-Based Featurization and Classification," U.S. Application No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing," U.S. Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," and U.S. Application No. 16 / 723,716, entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples," all of which are incorporated by reference in their entireties.

[0230] The analysis system can further restrict the CpG sites considered for use in the cancer classifier. For each CpG site in the initial set of CpG sites, the analysis system calculates information gain based on the feature vector of the training sample (step 530). From step 520, each training sample has a feature vector that includes anomaly scores of all CpG sites in the initial set of CpG sites. However, some CpG sites in the initial set of CpG sites may not be as informative as other CpG sites in distinguishing cancer types, or may overlap with other CpG sites.

[0231] In one embodiment, the analysis system calculates an information gain for each cancer type and for each CpG site in the initial set and determines whether to include that CpG site in the classifier (step 530). The information gain is calculated for a training sample with a given cancer type compared to all other samples. For example, two random variables "anomalous fragment" ("AF") and "cancer type" ("CT") are used. In one embodiment, AF is a binary variable indicating whether there is an anomalous fragment overlapping a given CpG site in a given sample, as determined for the anomaly score / feature vector above. CT is a random variable indicating whether the cancer is of a particular type. The analysis system calculates the mutual information for CT given AF. That is, it calculates how many bits of information about the type of cancer is obtained if we know whether there is an anomalous fragment overlapping a particular CpG site. In practice, for a first cancer type, the analysis system calculates the pairwise mutual information gain for each other cancer type and also sums the mutual information gains over all other cancer types.

[0232] For a given cancer type, the analysis system can use this information to rank CpG sites based on how cancer-specific they are. This procedure can be repeated for all cancer types under consideration. If a particular region is commonly aberrantly methylated in training samples of one cancer but not in training samples of other cancer types or in healthy training samples, then CpG sites that overlap with those aberrant fragments may have high information gain for a cancer type. The ranked CpG sites for each cancer type can be greedily added (selected) to a set of CpG sites selected based on their rank for use in the cancer classifier (step 540).

[0233] In additional embodiments, the analysis system can consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion is that the selected CpG sites have a threshold or greater separation from other selected CpG sites. For example, the selected CpG sites are separated from other selected CpG sites by a threshold or greater number of base pairs (e.g., 100 base pairs), such that CpG sites that are within the threshold separation are not both selected for consideration in the cancer classifier.

[0234] In one embodiment, the analysis system may optionally modify the feature vector of the training samples according to the set of CpG sites selected from the initial set (step 550). For example, the analysis system may truncate the feature vector to remove abnormality scores corresponding to CpG sites that are not included in the set of selected CpG sites.

[0235] The analysis system can train a cancer classifier in any of a number of ways using the feature vector of the training samples. The feature vector can correspond to the initial set of CpG sites from step 520, or a selected set of CpG sites from step 550. In one embodiment, the analysis system trains a binary cancer classifier (step 560) to distinguish between cancer and non-cancer based on the feature vector of the training samples. Thus, the analysis system uses training samples that include both non-cancer samples from healthy individuals and cancer samples from subjects. Each training sample can have one of two labels: "cancer" or "non-cancer". In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.

[0236] In another embodiment, the analysis system trains a multi-class cancer classifier (step 570) to distinguish between many cancer types (also referred to as tissue of origin (TOO) labels). The cancer types can include one or more cancers and can also include non-cancer types (including any additional other diseases or genetic disorders, etc.). To do so, the analysis system can use a cancer type cohort and may or may not include a non-cancer type cohort. In this multi-cancer embodiment, the cancer classifier is trained to determine a cancer prediction (more specifically, a TOO prediction) consisting of a predictive value for each of the cancer types classified. The predictive value may correspond to the likelihood that a given training sample (and the test sample under inference) has each cancer type. In one embodiment, the predictive value is scored from 0 to 100, with the cumulative predictive value equal to 100. For example, the cancer classifier returns a cancer prediction that includes predictive values ​​for breast cancer, lung cancer, and non-cancer. For example, the classifier can return a cancer prediction that the test sample has a 65% chance of breast cancer, a 25% chance of lung cancer, and a 10% chance of non-cancer. The analysis system may further evaluate the predictive values ​​to generate one or more predictions of the presence of cancer in the sample, which may also be referred to as TOO predictions indicating one or more TOO labels, e.g., a first TOO label having the highest predictive value, a second TOO label having the second highest predictive value, etc. Continuing with the example above, given the proportions, in this example the system may determine that the sample is afflicted with breast cancer since breast cancer has the highest probability.

[0237] In any of the embodiments, the analysis system trains a cancer classifier by inputting a set of training samples having feature vectors into the cancer classifier and adjusting classification parameters so that the function of the classifier accurately associates the training feature vectors with corresponding labels. The analysis system can group the training samples into one or more sets of training samples for iterative batch training of the cancer classifier. After inputting all sets of training samples including the training feature vectors and adjusting the classification parameters, the cancer classifier is sufficiently trained to label test samples according to the feature vectors within a certain error range. The analysis system can train the cancer classifier according to any one of a number of methods. As an example, a binary cancer classifier may be an L2-regularized logistic regression classifier trained using a logarithmic loss function. As another example, a multi-cancer classifier may be a multinomial logistic regression. In practice, any type of cancer classifier can be trained using other techniques. These techniques may include, among many others, machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autoencoder models, multi-layer neural networks, etc.

[0238] The classifier can include a logistic regression algorithm, a neural network algorithm, a support vector machine algorithm, a naive bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm.

[0239] <III.C. Development of Cancer Classifier> During use of the cancer classifier, the analysis system may obtain a test sample from a subject with an unknown type of cancer. The analysis system may process the test sample, which is comprised of DNA molecules, in any combination of processes 300, 400, and 420 to obtain a set of abnormal fragments. The analysis system may determine a test feature vector for use in the cancer classifier according to similar principles discussed in process 500. The analysis system may calculate an abnormality score for each CpG site in the plurality of CpG sites used by the cancer classifier. For example, the cancer classifier may receive as input a feature vector that includes an abnormality score for 1,000 selected CpG sites. Thus, the analysis system may determine a test feature vector that includes an abnormality score for the 1,000 selected CpG sites based on the set of abnormal fragments. The analysis system may calculate the abnormality score in the same manner as the training sample. In some embodiments, the analysis system defines the abnormality score as a binary score based on whether or not there is a hypermethylated or hypomethylated fragment encompassing the CpG site in the set of abnormal fragments.

[0240] The analysis system can then input the test feature vector into a cancer classifier. The cancer classifier function can then generate a cancer prediction based on the classification parameters trained in process 500 and the test feature vector. In a first embodiment, the cancer prediction is binary and can be selected from the group consisting of "cancer" or "non-cancer", and in a second embodiment, the cancer prediction is selected from the group consisting of a number of cancer types and "non-cancer". In additional embodiments, the cancer prediction has a predictive value for each of a number of cancer types. Furthermore, the analysis system can determine that the test sample is most likely to be one of the cancer types. Following the above example where the cancer prediction for the test sample is 65% chance of breast cancer, 25% chance of lung cancer, and 10% chance of non-cancer, the analysis system can determine that the test sample is most likely to be affected by breast cancer. In another example, if the cancer prediction is binary, 60% chance of non-cancer, and 40% chance of cancer, the analysis system determines that the test sample is most likely to be non-cancer. In additional embodiments, the cancer prediction with the highest likelihood may also be compared to a threshold (e.g., 40%, 50%, 60%, 70%) for calling the subject having that cancer type. If the cancer prediction with the highest likelihood does not exceed that threshold, the analysis system may return an inconclusive result.

[0241] In additional embodiments, the analysis system chains the cancer classifier trained in step 560 of process 500 with step 570 of process 500 or with another cancer classifier trained in process 500. The analysis system can input the test feature vector to a cancer classifier trained as a binary classifier in step 560 of process 500. The analysis system can receive an output of a cancer prediction. The cancer prediction can be a binary value for whether the subject is likely to have cancer or likely not to have cancer. In other embodiments, the cancer prediction includes a predictive value describing the likelihood of cancer and the likelihood of being non-cancer. For example, the cancer prediction has a cancer predictive value of 85% and a non-cancer predictive value of 15%. The analysis system can determine that the subject is likely to have cancer. If the analysis system determines that the subject is likely to have cancer, the analysis system can input the test feature vector to a multi-class cancer classifier trained to distinguish between different cancer types. The multi-class cancer classifier can receive the test feature vector and return a cancer prediction for the cancer type for the multiple cancer types. For example, the multi-class cancer classifier provides a cancer prediction that specifies that the subject is most likely to have ovarian cancer. In another embodiment, the multi-class cancer classifier provides a predictive value for each cancer type of a plurality of cancer types. For example, the cancer predictive values ​​can include a breast cancer type predictive value of 40%, a colon cancer type predictive value of 15%, and a liver cancer predictive value of 45%.

[0242] According to a generalized embodiment of binary cancer classification, the analysis system can determine a cancer score for the test sample based on the sequencing data of the test sample (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analysis system can compare the cancer score of the test sample against a binary threshold cutoff for predicting whether the sample is likely to be cancerous. The binary threshold cutoff can be adjusted using TOO thresholding based on one or more TOO subtype classes. The analysis system can further generate a feature vector for the test sample for use in a multi-class cancer classifier to determine a cancer prediction indicative of one or more likely cancer types.

[0243] The classifier can be used to determine the disease state of a subject, for example, a subject with an unknown disease state. The method can include obtaining a test genomic data construct (e.g., single-time point test data) in electronic form, the test genomic data construct comprising a value for each genomic feature in a plurality of genomic features of a corresponding plurality of nucleic acid fragments in a biological sample obtained from the subject. The method can then include applying the test genomic data construct to the test classifier, thereby determining the state of the disease state in the subject. The subject may not have been previously diagnosed with a disease state.

[0244] The classifier may be a temporal classifier that uses at least (i) a first test genomic data construct generated from a first biological sample obtained from the subject at a first time point, and (ii) a second test genomic data construct generated from a second biological sample obtained from the subject at a second time point.

[0245] The trained classifier can be used to determine the disease state of a subject, for example, a subject whose disease state is unknown. In this case, the method can include obtaining, in electronic form, a test time series dataset for the subject, where the test time series dataset includes, for each of a plurality of time points, corresponding test genotype data constructs including values for a plurality of genotype characteristics of a corresponding plurality of nucleic acid fragments in a corresponding biological sample obtained from the subject at each of the respective time points, and, for each pair of consecutive time points among the plurality of time points, a representation of the length of time between each pair of consecutive time points. The method can then include applying the test genotype data construct to a test classifier, thereby determining the state of the disease state in the subject. The subject may not have been previously diagnosed with a disease state.

[0246] <IV. Applications> In some embodiments, the methods, analysis systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor the progression or recurrence of cancer, monitor treatment response or efficacy, determine or monitor the presence of minimal residual disease (MRD), or any combination thereof. For example, as described herein, a classifier can be used to generate a probability score (e.g., 0-100) that describes the likelihood that a test feature vector is from a subject having cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether the subject has cancer. In other embodiments, the likelihood score or probability score can be evaluated at a plurality of different time points (e.g., before or after treatment) to monitor the progression of the disease or the efficacy of the treatment (e.g., treatment effect). In still other embodiments, the likelihood score or probability score can be used to make clinical decisions (e.g., cancer diagnosis, treatment selection, evaluation of treatment effect, etc.) or to influence clinical decisions. For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe an appropriate treatment.

[0247] <IV.A. Early Detection of Cancer> In some embodiments, the methods and / or classifiers of the present invention are used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (e.g., described above in Section III and discussed in Section V) can be used to determine a cancer prediction that describes the likelihood that a test feature vector is from a subject having cancer.

[0248] In one embodiment, the cancer prediction is the likelihood that a test sample has cancer (e.g., scored from 0 to 100) (i.e., binary classification). Thus, the analysis system can determine a threshold for determining whether a subject has cancer. For example, a cancer prediction of 60 or greater can indicate that the subject has cancer. In still other embodiments, a cancer prediction of 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater indicates that the subject has cancer. In other embodiments, the cancer prediction can indicate the severity of the disease. For example, a cancer prediction of 80 may indicate a more severe form, or a later stage of cancer, compared to a cancer prediction of less than 80 (e.g., a probability score of 70). Similarly, an increase in the cancer prediction value over time (e.g., determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in the cancer prediction value over time may indicate treatment success.

[0249] In another embodiment, the cancer prediction is composed of multiple predictive values ​​and classified (i.e., multi-class classification), with each of multiple cancer types having a predictive value (e.g., scored from 0 to 100). The predictive values ​​may correspond to the likelihood that a given training sample (and during inference, the training sample) has each of the cancer types. The analysis system may identify the cancer type with the highest predictive value and indicate that the subject is more likely to have that cancer type. In other embodiments, the analysis system further compares the highest predictive value to a threshold (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to determine that the subject is more likely to have that cancer type. In other embodiments, the predictive value may also indicate disease severity. For example, a predictive value above 80 may indicate a more severe form of cancer, or a later stage, compared to a predictive value of 60. Similarly, an increase in predictive value over time (e.g., as determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in predictive value over time may indicate successful treatment.

[0250] According to aspects of the invention, the methods and systems of the invention can be trained to detect or classify multiple cancer indications. For example, the methods, systems and classifiers of the invention can be used to detect the presence of 1 or more, 2 or more, 3 or more, 5 or more, 10 or more, 15 or more, or 20 or more different types of cancer.

[0251] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include carcinomas, lymphomas, blastomas, sarcomas, and leukemias or lymphoid malignancies. More specific examples of such cancers include squamous cell carcinomas (e.g., epithelial squamous cell carcinomas), skin cancers, malignant melanomas, lung cancers including small cell lung cancers, non-small cell lung cancers ("NSCLC"), lung cancers including lung adenocarcinomas and lung squamous cell carcinomas, cancers of the peritoneum, gastric or gastrointestinal cancers including gastrointestinal cancers, pancreatic cancers (e.g., pancreatic ductal adenocarcinomas), cervical cancers, ovarian cancers (e.g., high-grade serous ovarian cancers), liver cancers (e.g., hepatocellular carcinoma (HCC)), hepatocellular carcinomas, hepatomas, bladder cancers (e.g., urothelial bladder cancers), testicular cancers, and the like. (germ cell tumor) cancer, breast cancer (e.g., HER2-positive breast cancer, HER2-negative breast cancer, triple-negative breast cancer), brain tumors (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney or renal cancer (e.g., renal cell carcinoma, nephroblastoma or Wilms tumor), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal cancer (NPC). Additional examples of cancer include, but are not limited to, retinoblastoma, theca cell carcinoma, androgenic tumor, non-Hodgkin's lymphoma (NHL), multiple myeloma, and hematological malignancies, including but not limited to acute hematological malignancies, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteogenic sarcoma, leiomyosarcoma, and urinary tract cancer.

[0252] In some embodiments, the cancer is one or more of anal cancer, bladder cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, uterine cancer, or any combination thereof.

[0253] In some embodiments, one or more cancers can be anal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, pancreatic cancer, and "high-signal" cancers such as lymphoma and multiple myeloma (defined as cancers with a 5-year cancer-specific mortality rate exceeding 50%). High-signal cancers tend to be more aggressive, and generally, the concentration of cell-free nucleic acids in test samples taken from patients is above average.

[0254] <IV.B. Cancer and Treatment Monitoring> In some embodiments, cancer prediction can be evaluated at multiple different time points (e.g., before and after treatment or otherwise) to monitor the progression of the disease or the effectiveness of treatment (e.g., treatment effect). For example, the present invention includes a method comprising obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point and determining a first cancer prediction therefrom (as described herein), and obtaining a second test sample (e.g., a second plasma cfDNA sample) from the cancer patient at a second time point and determining a second cancer prediction therefrom (as described herein).

[0255] In certain embodiments, the first time point is before cancer treatment (e.g., before resection surgery or therapeutic intervention) and the second time point is after cancer treatment (e.g., after resection surgery or therapeutic intervention), and the classifier is utilized to monitor the effectiveness of the treatment. For example, if the cancer predictor value at the second time is decreased compared to the cancer predictor value at the first time, the treatment is deemed successful. However, if the cancer predictor value at the second time is increased compared to the cancer predictor value at the first time, the treatment is deemed unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after cancer treatment (e.g., after resection surgery or therapeutic intervention). In still other embodiments, cfDNA samples can be obtained and analyzed from a cancer patient at the first and second time points. For example, to monitor cancer progression, to determine whether the cancer is in remission (e.g., after treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effectiveness of a treatment (e.g., therapeutic).

[0256] One skilled in the art will readily understand that in order to monitor the state of cancer in a patient, test samples can be obtained from a cancer patient over any desired set of time points and analyzed according to the method of the present invention. In some embodiments, the first and second time points are a time amount from about 15 minutes to about 30 years, for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, for example about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 50 days, or for example about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or about 12 months, or for example about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or about 30 years. In other embodiments, test samples can be obtained from the patient at least once every 5 months, at least once every 6 months, at least once every year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.

[0257] <IV.C. Treatment> In yet another embodiment, the cancer prediction can be used to make clinical decisions (e.g., cancer diagnosis, treatment selection, evaluation of treatment efficacy, etc.) or to influence clinical decisions. For example, in one embodiment, if the cancer prediction (e.g., for cancer or for a specific cancer type) exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., resection surgery, radiation therapy, chemotherapy, and / or immunotherapy).

[0258] The classifier (described herein) can be used to determine a cancer prediction that the sample feature vector is from a subject with cancer. In one embodiment, if the cancer prediction value exceeds a threshold, an appropriate treatment (e.g., resection surgery or treatment) is prescribed. For example, in one embodiment, if the cancer prediction value is greater than or equal to 60, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction value is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more appropriate treatments are prescribed. In other embodiments, the cancer prediction can indicate the severity of the disease. An appropriate treatment matching the severity of the disease can then be prescribed.

[0259] In some embodiments, the treatment is one or more cancer therapeutics selected from the group consisting of chemotherapeutic agents, targeted cancer therapeutics, differentiating therapeutics, hormonal therapeutics, and immunotherapeutic agents. For example, the treatment can be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disrupting agents (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutics selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiating therapeutics comprising retinoids such as tretinoin, alitretinoin, bexarotene. In some embodiments, the treatment is one or more hormonal therapeutics selected from the group consisting of antiestrogen agents, aromatase inhibitors, luteinizing hormone agents, estrogen agents, antiandrogen agents, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group consisting of monoclonal antibody therapies such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapies and adjuvants such as BCG, interleukin-2 (IL-2) and interferon-alpha, and immunomodulatory drugs such as thalidomide and lenalidomide (REVLIMID). Selecting an appropriate cancer therapeutic based on characteristics such as tumor type, cancer stage, prior cancer treatment or exposure to therapeutic agents, and other characteristics of the cancer is within the ability of a skilled physician or oncologist.

[0260] <V. Kit Implementation> Also disclosed herein are kits for carrying out the above-mentioned methods, including the methods related to cancer classifiers. The kits may include one or more collection containers for collecting samples from individuals containing genetic material. The samples may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such kits may include reagents for isolating nucleic acids from the samples. The reagents may further include reagents for sequencing the nucleic acids, including buffers and detection agents. In one or more embodiments, the kits may include one or more sequencing panels that include probes for targeting specific genomic regions, specific mutations, specific genetic mutations, or combinations thereof. In one or more embodiments, the kits include at least one panel that includes probes targeting contaminations, for example, from Table 2, Table 4, or some combination thereof. In other embodiments, the samples collected via the kits are provided to a sequencing laboratory that may use the sequencing panels to sequence the nucleic acids in the samples.

[0261] The kit may further include instructions for use of the reagents included in the kit. For example, the kit may include instructions for taking a sample and extracting nucleic acid from a test sample. Examples of instructions may include the order of adding reagents, the centrifugation speed to use to isolate nucleic acid from a test sample, a method for amplifying nucleic acid, a method for sequencing nucleic acid, or a combination thereof. The instructions may further reveal how to operate a computing device as the analysis system 200 to perform any of the steps of the described methods.

[0262] In addition to the above components, the kit can include a computer-readable storage medium storing computer software for performing the various methods described throughout this disclosure. One form in which these instructions may exist is information printed on a suitable medium or substrate, such as information printed on a piece of paper or pieces of paper, information printed in the packaging of the kit, in a package insert. As yet another means, it is a computer-readable medium, such as a diskette, CD, hard drive, network data storage device, on which instructions are stored in the form of computer code. As yet another means, there is a website address accessible via the Internet to the information of a deleted site.

[0263] [VI. Example Results] <VI.A. Sample Collection and Processing> The CCGA (NCT02889978), which is the study design and sample, is a multi-site collaborative study, case-control study, and observational study with long-term follow-up. Biospecimens from approximately 15,000 participants were collected from 342 sites. The samples were divided into a training set (1,785 cases) and a test set (1,015 cases), and the samples were selected so that the distribution of cancer and non-cancer types between sites in each cohort would be as pre-specified, and the cancer and non-cancer samples were frequency-matched by gender and age.

[0264] Whole-genome bisulfite sequencing was employed to isolate cfDNA from plasma and analyze the cfDNA by whole-genome bisulfite sequencing (WGBS; 30x depth). cfDNA was extracted from two plasma tubes per patient (up to 10 ml total volume) using a modified QIAamp Circulating Nucleic Acid Kit (Qiagen, Germantown, MD). Up to 75 ng of plasma cfDNA was converted to bisulfite using the EZ-96 DNA Methylation Kit (Zymo Research, D5003). The converted cfDNA was used to prepare dual-indexed sequencing libraries by using the Accel-NGS Methyl-Seq DNA Library Preparation Kit (Swift BioSciences, Ann Arbor, MI), and constructed libraries were quantified using the KAPA Library Quantification Kit for Illumina Platforms (Kapa Biosystems, Wilmington, MA). The four libraries were pooled together with 10% PhiX v3 library (Illumina, FC-110-3001) and clustered on an Illumina NovaSeq 7000 S2 flow cell followed by 150bp paired-end sequencing (30x).

[0265] For each sample, the WGBS fragment set was reduced to a small subset of fragments with abnormal methylation patterns. Furthermore, hypermethylated or hypomethylated cfDNA fragments were selected. cfDNA fragments with abnormal methylation patterns and being hypermethylated or hypomethylated, i.e., having UFXM, were selected. Fragments that occur frequently in non-cancerous individuals or fragments with unstable methylation are unlikely to produce highly discriminative features for classifying the cancer state. Therefore, an independent reference set (i.e., reference genome) consisting of 108 non-smokers without cancer (age: 58 ± 14 years, 79 females [73%]) obtained from the CCGA study was used to create a statistical model and the data structure of typical fragments. These samples were used for training a Markov chain model (order 3) that estimates the likelihood of a given sequence of CpG methylation states within a fragment, as described above in Section II.C. This model was demonstrated to be calibrated within the normal fragment range (p-value > 0.001) and was used to reject fragments with a p-value > 0.001 from the Markov model as being insufficiently abnormal.

[0266] As described above, in a further data reduction step, only fragments with at least 5 CpGs covered and an average methylation of 0.9 or more (hypermethylated) or less than 0.1 (hypomethylated) were selected. By this procedure, the median (range) of UFXM fragments for participants without cancer onset in the training was 2,800 (1,500 - 12,000), and the median (range) of UFXM fragments for participants with cancer onset in the training was 3,000 (1,200 - 420,000). Since this data reduction procedure uses only reference set data, this step needed to be applied only once to each sample.

[0267] <VI.B. Contamination Detection Results> As the first group of result examples, to examine the performance of the contamination marker probes in Tables 2 and 4, multiple batches of plasma cfDNA samples and gDNA samples from different internal studies were used for preliminary analysis, regardless of the presence or absence of intentional contamination, before further efforts were put into the development of the codebase.

[0268] FIG. 7 shows the distribution of the percentage of the number of fragments classified as contaminant relative to the total number of fragments called reference or alternative for each contaminant marker called homozygous for a given sample according to the first set of example results. The analytical system utilized the contaminant marker probes of Tables 2 and 4 to identify contaminant fragments in the samples. The percentage of contaminant fragments was calculated as the ratio of identified contaminant fragments to the total number of fragments in each sample that overlap with the contaminant marker probe. As shown in graph 700, the percentage of contaminant fragments for each sample was plotted with the Y-axis showing the percentage of contaminant fragments. Although there were a few samples that were above 0.001, the majority of samples were below 0.001. In the box plot of graph 700, the lower quartile, median, and upper quartile are all well below 0.0005. Graph 710 shows the same results, plotting the percentage of contaminant fragments on the X-axis and the number of samples on the Y-axis. See FIG. 17 for a corresponding diagram of the second set of example results.

[0269] FIG. 8 is a scatter plot of allele frequencies of contaminant markers according to the first example result. The analysis system calculated the allele frequencies of the contaminant markers using samples from the CCGA study and its follow-up study. The analysis system also obtained population allele fractions from the genome databases Genomes Project and gnomAD (graph 1000). Each contaminant marker is plotted with the x-axis showing the population allele fraction and the y-axis showing the allele fraction calculated from the samples. For a corresponding diagram of the second group of example results, see FIG. 14.

[0270] FIG. 9 shows two graphs showing the genotypes and zygosity of SNP contaminant markers (graph 900) and indel contaminant markers (graph 910) according to the first group of exemplary results. In the graph 900 for multiple SNP contaminant markers, the three box plots on the left represent the proportional breakdown of zygosity of the first set of samples (reference homozygote, alternative homozygote, and heterozygote), and the three box plots on the right represent the proportional breakdown of zygosity of the second set of samples (reference homozygote, alternative homozygote, and heterozygote). Similarly, the graph 910 for indel contaminant markers also represents the proportional breakdown of zygosity of indel contaminant markers divided between two sets of samples. See FIG. 15 and FIG. 16.

[0271] FIG. 10 is a graph showing the percentage of contaminant markers for a given sample that are homozygous and overlap enough fragments to be considered useful for estimating contamination of that sample according to the first set of example results. The analysis system determined the zygosity of the contaminant markers for each sample. The analysis system then calculated the percentage of contaminant markers that were found to be homozygous in the sample. The more contaminant markers found to be homozygous in a sample, the higher the contamination detection power. Furthermore, these homozygous sites go through various quality check criteria, such as overlapping enough sample fragments to be considered useful for estimating contamination of the sample, to obtain the final percentage of useful sites for a given sample. The percentages were specifically plotted for both multiple SNP site contaminant markers, indel site contaminant markers, multiple SNP site contaminant markers and indel site contaminant markers. In the first box plot 1000, the median percentage was about 0.525. In the second box plot 1010 for multiple SNP site contaminant markers, the median percentage was about 0.575. In the third box plot 1020, for indel site contaminant markers. The median value was about 0.475. See Figures 15 and 13.

[0272] 11A and 11B show graphs of estimated contamination level detection for different batches of samples according to the first set of example results. Graph 1100 shows the distribution of estimated sample contamination levels from a "Derisking" batch of samples. Graph 1110 shows the distribution of estimated sample contamination levels from a "doppler_prelim_test" batch of samples. Graph 1120 shows the distribution of estimated sample contamination levels from a "Hyb1_SOP_12plex" batch of samples. Graph 1130 shows the distribution of estimated sample contamination levels from a "MRD" batch of samples. Graph 1140 shows the distribution of estimated sample contamination levels from a "cfDNA titration" batch of samples. Graph 1150 shows the distribution of estimated sample contamination levels from a "gDNA titration" batch of samples. Graph 1160 shows the distribution of estimated sample contamination levels from the "downsampled gDNA titration (v1.0 chemistry)" batch of samples, and graph 1170 shows the distribution of estimated sample contamination levels from the "downsampled gDNA titration (v1.5 chemistry)" batch samples. See Figures 20 and 21 for corresponding examples of results for the second group.

[0273] As a second set of example results, to examine the performance of the contamination marker probes in Tables 2 and 4, a cohort of plasma cfDNA samples from 84 single individuals without intentional contamination was used in the final analysis using the revised and matured codebase.

[0274] FIG. 12 shows the distribution of the number of unique cfDNA fragments obtained for each contaminant marker listed in Tables 2 and 4, tabulated across the 84 samples in the experiment. The x-axis shows the number of unique cfDNA fragment molecules overlapping with probes designed to target the contaminant marker, tabulated across all 84 samples, and the y-axis shows the percentage of contaminant markers that fall into the bins indicated by the x-axis.

[0275] Fragment markers as listed in Tables 2 and 4 were selected without steps 220 and 250 of the analysis system. Markers not in Hardy-Weinberg equilibrium were filtered based on the population database frequencies of the reference and alternative haplotypes. In this case, the values ​​of the Hardy-Weinberg term (p2+2pq+q2) were in the range [0.9, 1.5] (a bias towards values ​​higher than 1 is expected in non-random mating conditions of the population). Of the 1000 markers in Tables 2 and 4, 4 markers failed this condition.

[0276] Each contaminating marker underwent further quality checks (QC) for individual samples for cfDNA fragments overlapping with that marker's probe. Markers that did not pass the quality check criteria for a particular sample were subsequently discarded for that sample for any analysis relying on genomic variant calls for that particular marker, including estimation of contamination fraction.

[0277] For each pair of contaminating markers and one sample from the 84 samples of the experiment, a minimum number of unique cfDNA fragment molecules overlapping the position range of the genomic variant, in this case 20, that can be used to reliably call the haplotypes and alleles of multiple SNP markers and indel markers as reference or alternative, respectively, can be used as reference or alternative markers for more than a certain percentage (e.g., 0.7 for multiple SNP markers and 0.5 for indel markers). The latter condition is extended to check that the ratio of the number of cfDNA fragments called reference haplotypes to the number of cfDNA fragments called alternative haplotypes should not be highly improbable (e.g., p-value < 10-5) according to the theoretical binomial probability distribution expected for the marker and sample pair (average of 0.5 for markers called heterozygous, or average depending on the total contamination rate of all markers in the sample called homozygous).

[0278] Figure 13 shows the results of this quality check process. Based on the criteria mentioned above, 851 markers passed for all 84 samples, 68 markers passed for 83 samples, and 20 markers passed for 82 samples. Seven markers failed for all 84 samples (including four that did not meet the Hardy-Weinberg equilibrium passing criteria). The failure rates of the remaining markers varied between these two extremes.

[0279] Figures 14-16 show the characteristics of the genomic variants underlying the contaminating markers in Tables 2 and 4 for those contaminating markers that passed the quality check criteria for at least 82 out of 84 samples. Figure 14 shows a scatter plot of the allele frequencies in the range [0.3, 0.7], which is the allele frequency range for selecting these contaminating markers. Figure 15 shows a scatter plot of the frequency at which the markers are called heterozygous. Figure 16 shows a scatter plot of the values ​​of the Hardy-Weinberg term. The x-axis of these graphs represents values ​​obtained or calculated from data available in the population database 1000 Genomes (for multiple SNP markers) or gnomAD (for indel markers), and the y-axis represents values ​​calculated from the 84 cfDNA samples in our experiment. All points were expected to lie with some noise around the y=x line (represented by a dashed line in the graphs).

[0280] To estimate the contamination rate in a set of cfDNA fragments, we first identified cfDNA fragments originating from contaminating sources. cfDNA fragments that were referred to as variants opposite homozygous variants, referred to as markers, were classified as contaminating fragments. The subset of fragments in which fragments did not overlap contaminating marker sites or overlapped contaminating marker sites that were genotyped as heterozygous were ignored for the purposes of classifying contaminating fragments and estimating the contamination rate. The ability to estimate the contamination rate was contingent on finding at least one contaminating fragment. Because contaminating fragments could not be fractionated, the lower limit of detectable contaminating fragments was inversely proportional to the total number of fragments considered.

[0281] Figure 17 shows the distribution of the percentage of the number of fragments classified as contaminants relative to the total number of fragments called reference or alternative for each contaminant marker that was called homozygous for a given specimen, tabulated for each marker across all 84 specimens. Graph 1710 shows a jitter scatter plot and box plots of these values ​​showing the extremes and quartiles of the distribution. The arithmetic mean of these values ​​is 2x10 -4 Values ​​below the detection limit are suppressed to 0, resulting in a small bump at 0. Graph 1720 shows a smoothed density line of the same data (solid line), and a smoothed density line of the same data with binomial parameter p=2x10 -4 The dashed line shows the simulated values ​​obtained by using a binomial fit to the data, where sample size is the number of cfDNA fragments for a particular contaminant marker. Overall, there is a good fit, although there is some deviation around 0 and the mode.

[0282] We developed the concept of contaminant fragments to estimate the contamination rate of a given cfDNA sample by modeling the underlying process of cfDNA contamination from external sources. Considering only markers called as homozygous for the sample, if the contamination rate is cf and the population allele frequency of the variant opposite the homozygous called for the sample is afi, then the probability of observing a contaminant fragment for a marker is cfxafi. If the total number of fragments at this marker is ni, then the expected number of contaminant fragments at this marker is cfxafixni. Using the linearity property of expectation in probability theory, we sum these expected numbers for all markers to obtain the formula for the expected total number of contaminant fragments across all markers, nc=Σ i (cfxaf i xn i ) can be obtained by rearranging all known variables to bring them to one side, and the estimate of the contamination fraction is cf=nc / (Σ i (af i xn i )).

[0283] Figure 18 illustrates the application of the formula to estimate contamination rates for a hypothetical set of fragments. In this example, there are four multiple SNPs and two indel contaminating markers, of which only three multiple SNPs and one indel marker are called homozygous. Considering the identification of the three contaminating fragments, the total number of fragments at each marker, and the allele frequency at each marker for the opposite haplotype, an estimate of the contaminating fragments for the entire sample is obtained.

[0284] Figure 19 shows the results of applying the contamination fraction model to the simulated data. For each point, data were generated by sampling variant calls (homozygous reference, homozygous alternative, or heterozygous) for both the background and contaminant genomes based on their population frequencies for each contaminant marker in Tables 2 and 4. The counts of the reference and alternative haplotype fragments were then sampled from a binomial distribution whose size was the average number of cfDNA fragments observed for this marker among the 84 samples obtained from the experiment, with the probability a function of the two variant calls and the contamination fraction entered as a simulation parameter. To capture the variability of the estimates, each simulation was repeated 100 times. The x-axis of the graph represents the contamination fraction used as a parameter in the simulation, and the y-axis represents the contamination fraction value estimated by the model. The dashed lines represent y=x-line. Each box plot shows the range of estimated contamination fractions for 100 simulations performed with the specified contamination fraction parameters. The range shows that the median is very close to the expected value, with the interquartile range becoming smaller as the degree of contamination increases. These results validate the model described herein for estimating contamination rates.

[0285] Figure 20 shows the results of applying the contamination rate model to four titration pairs of cfDNA samples with different titration levels. The donor samples in the titration pairs are treated as contamination, and their titration levels are treated as contamination rates. To confirm that reliable estimated values of the contamination rate are obtained, each titration level in each group was also repeated multiple times according to the availability of the input material. The X-axis represents the titration level, and the Y-axis represents the estimated contamination rate for each replicate titrated at that level. The solid line shows the linear model fit to all observed values, and the dashed line shows the y = x line. From this graph, the linear model fit has approximately the same slope as the y = x line, but the intercepts of the two lines are different, and there is a constant unintentional background contamination of about 2.5x10 -4 It can be seen that there is. When examining the contamination fragments in these samples in detail, it became clear that this constant residual contamination is real and not due to technical noise such as variant calling (not shown here). These results further validate the model described herein for estimating the contamination rate.

[0286] Figure 21 shows the distribution of the estimated contamination rates of 84 cfDNA samples in the experiment. This graph shows a jitter scatter plot as well as a box-and-whisker plot of the estimated contamination rate values. The median is 1.8x10 -4 and the mean is 3.2x10 -4 . It should be noted that this mean value is different from the mean value of the proportion of contamination fragments obtained in the same sample set (2x10 -4 ) and is high, which is presumably because the contamination fragment model takes into account contamination fragments that could not be identified when the markers have the same haplotype in both genomes.

[0287] Figure 22 shows a high correlation and no significant systematic bias in the estimated values obtained for 84 samples in the experiment when considering multiple SNP markers and indel markers separately. This demonstrates that the performance of both types of contamination markers is equivalent.

[0288] <VII. Additional Considerations> The above detailed description of the embodiments refers to the accompanying drawings which show specific embodiments of the present disclosure. Other embodiments having different structure and operation do not depart from the scope of the present disclosure. Terms such as "the invention" are used in reference to specific examples of the many alternative aspects or embodiments of Applicant's invention described herein, and their use or absence is not intended to limit the scope of Applicant's invention or the claims.

[0289] An embodiment of the present invention may also relate to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes and / or may consist of a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory, tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions that may be coupled to a computer system bus. Furthermore, the computing systems referred to herein may include a single processor or may be structures employing multiple processor designs to increase computing power.

[0290] Any of the steps, operations, or processes described herein as being performed by the analysis system may be performed or implemented using one or more hardware or software modules of the apparatus, alone or in combination with other computing devices. In one embodiment, the software modules are implemented in a computer program product including a computer readable medium containing computer program code that may be executed by a computer processor to perform any or all of the processes, operations, or steps described.

Claims

1. 1. A method of predicting the presence of cancer in a test sample, comprising: obtaining a test sample, the test sample comprising a plurality of sequence reads for cell-free DNA (cfDNA) fragments in the test sample; identifying one or more contaminant markers from a plurality of contaminant markers for which the test sample has a homozygous haplotype; identifying as contaminating cfDNA fragments any cfDNA fragments in the test sample having a haplotype at one of the identified contaminating markers that is different from the homozygous haplotype for the respective contaminating marker; estimating the level of contamination based on any identified contaminating cfDNA fragments; and determining whether the contamination level is below a threshold level; and responsive to determining that the contamination level is at or below the threshold level, performing a cancer classification on the sequence reads of the cfDNA fragments in the test sample to generate a cancer prediction; A method comprising:

2. 10. The method of claim 1, wherein the plurality of contamination markers comprises a plurality of single nucleotide polymorphism (SNP) sites.

3. 3. The method of claim 2, wherein the plurality of SNP sites are within 10 base pairs (bps).

4. 3. The method of claim 2, wherein the plurality of SNP sites have a population haplotype frequency within the range of 45% to 55%.

5. 3. The method of claim 2, wherein the plurality of SNP sites excludes guanine-adenine polymorphisms and cytosine-thymine polymorphisms.

6. 3. The method of claim 2, wherein the haplotypes at each of the plurality of SNP sites are in Hardy-Weinberg equilibrium.

7. 3. The method of claim 2, wherein the plurality of contaminant markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of the plurality of SNP sites.

8. 3. The method of claim 2, wherein the plurality of contaminant markers comprises the plurality of SNP sites in Table 1.

9. 10. The method of claim 1, wherein the plurality of contaminating markers comprises insertion-deletion (indel) sites.

10. 10. The method of claim 9, wherein the indel site is 5 bps to 10 bps.

11. 10. The method of claim 9, wherein the indel sites have a population haplotype frequency in the range of 45% to 55%.

12. 10. The method of claim 9, wherein the haplotypes at each indel site are in Hardy-Weinberg equilibrium.

13. 10. The method of claim 9, wherein the plurality of contamination markers comprises at least 500, at least 1,000, at least 1,500, or at least 2,000 of the indel sites.

14. 10. The method of claim 9, wherein the plurality of contamination markers comprises the indel sites of Table 3.

15. 10. The method of claim 1, wherein each contaminant marker comprises a probe designed to target each haplotype of said contaminant marker.

16. 2. The method of claim 1, wherein the estimation of the contamination level is further based on one or more of the number of identified contaminating cfDNA fragments, the sequencing depth of the test sample, the number of the cfDNA fragments in the test sample, and the number of the contaminating markers.

17. 10. The method of claim 1, wherein in response to determining that the contamination level is equal to or greater than the threshold level, a classification of cancer is foregone.

18. 10. The method of claim 1, wherein the cancer prediction comprises a binary prediction between cancer and non-cancer.

19. 10. The method of claim 1, wherein the cancer prediction comprises a multi-class cancer prediction among multiple cancer types.

20. 10. The method of claim 1, wherein performing the cancer classification comprises: generating a test feature vector based on sequence reads of cfDNA fragments in the test sample; and inputting the test feature vector into a classification model to generate the cancer prediction for the test sample; A method comprising:

21. 21. The method of claim 20, wherein performing the cancer classification further comprises: filtering the initial set of cfDNA fragments of the test sample with p-value filtering to generate a set of aberrant fragments; the filtering comprises removing fragments from the initial set that are below a threshold p-value with respect to other fragments to generate the set of abnormal fragments; wherein the test feature vector is based on the sequence reads of the set of abnormal fragments. method.

22. 21. The method of claim 20, wherein the classification model is a machine learning model.

23. The method of claim 1, The method further comprising generating a notification indicating that the test sample is contaminated in response to determining that the contamination level is equal to or greater than the threshold level.

24. A non-transitory computer readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform the method of any one of claims 1 to 22.