Leukocyte contamination detected

A computer model for WBC contamination detection in cell-free DNA samples addresses the limitations of current cancer detection methods, enhancing accuracy and enabling comprehensive screening for multiple cancer types.

JP2026509730APending Publication Date: 2026-03-25GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current cancer detection methods are cancer type-specific, leading to low detection rates and high false-positive rates, and are hampered by leukocyte contamination in blood samples distorting cancer analysis.

Method used

A computer model for WBC contamination detection in cell-free DNA samples, which evaluates and corrects for leukocyte contamination, enabling comprehensive screening for multiple cancer types from a single sample using next-generation sequencing technologies.

Benefits of technology

Improves cancer detection accuracy by reducing false positives and ensuring sample quality, allowing for early and multi-type cancer screening from a single sample.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509730000001_ABST
    Figure 2026509730000001_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting leukocyte (WBC) contamination in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments. The computer-implemented method for WBC contamination detection aims to evaluate whether the sample is contaminated with WBC efflux DNA and can further determine the level of contamination.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims the benefits and priority of U.S. Provisional Patent Application No. 63 / 489,986 filed on 13 March 2023, U.S. Provisional Patent Application No. 63 / 495,766 filed on 12 April 2023, and U.S. Provisional Patent Application No. 63 / 518,881 filed on 11 August 2023, all of which are incorporated by reference. [Background technology]

[0002] Cancer is a leading cause of death worldwide. Cancer mortality is exacerbated by the fact that cancer is typically detected in its later stages, limiting the effectiveness of treatment options for long-term survival. Current detection methods are generally cancer type-specific; that is, each cancer type is screened individually. Each individual screening process is tailored to the specific cancer type. For example, mammography scans are used to detect breast cancer, while colonoscopy or stool tests are useful for detecting colorectal cancer. Each diverse screening method is not applicable to other cancer types. For example, to screen an individual for three different possible cancer types, healthcare providers must perform or instruct them to perform three different screening processes. Each of these screening processes may involve a combination of invasive and / or non-invasive procedures to identify tumor growth, take biopsies of the growths, and perform analysis on the tissue biopsies.

[0003] Furthermore, this screening method is hampered by low detection rates or high false-positive rates. Low detection rates often fail to detect early-stage cancers because they are in their very early stages. High positive rates lead to misdiagnosis of cancer in individuals who do not have cancer. As a result, most screening tests are used to examine individuals at high risk of developing the cancer they are screening for, and are only practical when their ability to detect cancer in the general population is limited.

[0004] Novel research is linked to abnormal DNA methylation in many disease processes, including cancer. DNA methylation plays a role in regulating gene expression. Therefore, abnormal DNA methylation can disrupt normal gene expression pathways, potentially leading to cancer or other diseases. For example, specific patterns of differentially methylated regions may be useful as molecular markers for various pathological conditions. Nevertheless, even such models face several challenges. Early cancer detection is particularly difficult due to the extremely small ratio of tumor cells to non-cancerous cells in the subject. This extremely small ratio can be as low as 1:1,000, 1:10,000, or even 1:100,000. This creates the challenge of detecting small amounts of cancer signaling within healthy signals. In addition, DNA can be excreted by blood cells, which may contain age-related genetic mutations that often resemble cancerous abnormal methylation. These highly informative methylated fragments excreted from blood cells can often artificially inflate cancer signaling.

[0005] During cfDNA sample preparation, further challenges arise due to the contamination of blood samples with leukocytes (WBCs) from the plasma. WBC signals also generally contain high levels of genetic variation, stemming from the natural behavior of leukocytes. However, these WBC signals are generally predicted to be artificial cancer signals, thereby distorting the analysis of cancer classification using WBC-contaminated samples. [Overview of the project] [Problems that the invention aims to solve]

[0006] This disclosure relates to addressing the issues referenced above. The background art provided herein is intended to provide a general context for this disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims of this application and, by being included in this section, shall not be considered prior art or a proposal of prior art. [Means for solving the problem]

[0007] The present invention and disclosure described herein offer improvements to cancer detection and treatment, particularly improvements to contamination detection so that sample quality is reliably maintained from sampling to prediction through sequencing and analysis. The present invention describes one or more embodiments of contamination prediction models. The contamination prediction model performs analysis using sequencing data with one or more contamination markers to evaluate WBC contamination in a sample.

[0008] Corrective measures can be taken for samples deemed contaminated by WBC-extracted cfDNA. For example, to prevent skewed results, the use of contaminated samples in downstream analyses can be avoided. In addition, contaminated samples may be physically discarded or otherwise labeled as unusable. New samples can also be taken from the individual. Furthermore, a contamination detection workflow can be used to evaluate the source of contamination. For example, a contamination detection workflow can evaluate the quality of the sample before and after steps in the sample processing workflow. If the quality of the sample is degraded, this indicates that the testing step contributed to the contamination. Alternatively, as another example, a contamination detection workflow may be used to evaluate sample quality from two parallel workflows (or two sequencing assays) to determine which sample quality is better maintained. This improved contamination detection ensures the accuracy of analytical predictions and improves the sample assay to avoid prediction reverts based on contaminated samples.

[0009] The present invention includes screening for cancer signals in a cell-free deoxyribonucleic acid (cfDNA) sample of interest. Such a cfDNA sample may contain thousands, tens of thousands, hundreds of thousands, millions, or more cfDNA fragments, thereby obtaining a similar order of sequence reads output by the sequencer, or multiples of such orders based on the sequencing depth of the sample. The length of each sequence read associated with the cfDNA fragment can vary, for example, up to 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 bp in length. These next-generation sequencing technologies significantly increase the amount of fragments that can be sequenced and analyzed, thereby enabling these models to identify even trace amounts of cancer signals in a sample. The present invention allows for screening of cancer or multiple cancer types from a single sample. This is an improvement over conventional screening methods tailored to specific cancer types by providing a single, comprehensive screening that can screen for various cancer types from a single cfDNA sample. For example, screening for different cancers generally involved separate screening processes that targeted the detection of tissue proliferation.

[0010] This invention implements a computer model for evaluating whether a sample is contaminated with WBC-shedded DNA. This process may be called WBC contamination detection. Generally, blood samples used for cancer detection are processed to separate plasma, a buffy coat containing WBCs, and solid red blood cells and other microparticles. Plasma generally contains cfDNA fragments shed from cells in the body. In some patients with cancer, tumor cells shed cfDNA fragments into the plasma. Therefore, cancer signals can be detected by analysis of cfDNA. However, leukocytes also tend to shed DNA fragments during their normal functioning. Such DNA shed by leukocytes, when analyzed together with cfDNA in plasma, can interfere with cancer detection and may lead to false-positive calls for cancer. The computer model implemented in WBC contamination detection aims to evaluate whether a sample is contaminated with WBC-shedded DNA and may further determine the level of contamination. Such screening ensures that the training data used to train downstream models is not distorted due to WBC contamination. Screening can also reduce false positive predictions by refraining from using contaminated samples for further analysis. If a contaminated sample is detected, further corrective actions can be taken to address the contamination. Removal of the source of contamination improves the physical assay process and sample handling.

[0011] The present invention implements a computer model for identifying and quantifying cancer signals. In one or more embodiments, the computer model may identify informative methylation fragments (also referred to as fragments having informative methylation patterns). From the informative methylation fragments, the computer model can be trained to identify features for use in cancer classification. The computer model may include a trained cancer classifier configured to input a feature vector generated based on the informative methylation fragments and output a cancer prediction based on the input feature vector. The cancer prediction can be a binary prediction and / or a multi-class prediction. The binary prediction can be the likelihood of the presence of cancer. The multi-class prediction can be the likelihood of a specific cancer type from a plurality of cancer types being evaluated. By training a cancer classifier that can screen among a plurality of cancer types, medical professionals can utilize a single comprehensive screening rather than a plurality of different screenings. In one or more embodiments, the cancer classifier is a machine learning model based on computer capabilities and is not substantially executable by human thinking. The cancer prediction can be used by a healthcare provider for diagnosis, prognosis, treatment adjustment, treatment evaluation, minimal residual disease detection, and the like.

[0012] A method for identifying a contamination marker for detecting contamination of white blood cells (WBC) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: obtaining a first set of cfDNA samples from a first cohort of subjects and a second set of WBC samples from the first cohort of subjects; determining, for each sample, the coverage of sequence reads overlapping each genomic locus in an initial set of genomic loci; determining, for each genomic locus, the ratio of coverage between the WBC sample and the cfDNA sample; and determining a set of features of the genomic loci as contamination markers where the ratio of coverage exceeds a threshold.

[0013] Clause 2. The method according to Clause 1 or any other clause dependent thereon, wherein each cfDNA sample consists essentially of sequence reads corresponding to cfDNA fragments, and each WBC sample consists essentially of sequence reads corresponding to DNA fragments discharged from WBCs.

[0014] Clause 3. The method according to Clause 1 or any other clause dependent thereon, wherein each subject in the first cohort is associated with at least one cfDNA sample and at least one WBC sample.

[0015] Clause 4. The method according to Clause 1 or any other clause dependent thereon, further comprising normalizing the coverage of sequence reads overlapping each genomic locus based on the sequencing depth of the sample for each sample.

[0016] Clause 5. The method according to Clause 1 or any other clause dependent thereon, wherein each genomic locus includes at least one CpG site.

[0017] Clause 6. The method according to Clause 1 or any other clause dependent thereon, further comprising pruning one or more genomic loci with highly variable coverage across samples to generate an initial set of genomic loci.

[0018] Clause 7. The method according to Clause 1 or any other clause dependent thereon, further comprising pruning one or more genomic loci with coverage that is non-normally distributed across samples.

[0019] Clause 8. For each genomic locus, the method according to Clause 3, comprising calculating the ratio of the coverage of the genomic locus from the WBC sample to the coverage of the genomic locus from the cfDNA sample for each subject in the first cohort to obtain the ratio of the coverage of each genomic locus, and determining the representative value of the ratio of the coverage of the genomic locus across the subjects in the first cohort as the ratio of the coverage of each genomic locus.

[0020] Clause 9. A method for training the detection of leukocyte (WBC) contamination in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: obtaining a set of cfDNA samples from a non-cancer subject, each cfDNA sample containing sequence reads corresponding to a cfDNA fragment; determining, for each sample, the coverage of sequence reads overlapping with each genomic locus in a set of genomic locus features as contamination markers; and generating a coverage distribution for each genomic locus in the set of genomic locus features based on the coverage of a first subset of cfDNA samples. A method comprising: determining the statistical likelihood of observing coverage at each genomic locus based on the coverage distribution of genomic loci for each cfDNA sample in a subset of two; generating contamination metrics by combining the statistical likelihoods of cfDNA samples between genomic loci for each cfDNA sample in a second subset; determining the distribution of contamination metrics for a second subset of cfDNA samples; and determining a contamination threshold based on the distribution of contamination metrics, wherein the contamination model includes the coverage distribution across feature sets of genomic loci, the distribution of contamination metrics, and the contamination threshold.

[0021] Clause 10. The method of Clause 9 or any other clause subordinate thereto, further comprising normalizing the coverage of sequence reads that overlap with each genomic locus based on the sequencing depth of the sample for each sample.

[0022] Clause 11. Each genomic locus comprises at least one CpG site, as described in Clause 9 or any other clause subordinate thereto.

[0023] Clause 12. A set of feature quantities for a genomic locus is identified in accordance with the method described in Clause 1 or any other clause subordinate thereto, as described in Clause 9 or any other clause subordinate thereto.

[0024] Clause 13. The statistical likelihood of observing coverage at each genomic locus is the z-score based on the mean and standard deviation defining the distribution of the genomic locus, as described in Clause 9 or any other clause subordinate thereto.

[0025] Clause 14. Contamination metrics are calculated by summing the absolute values ​​of the z-scores, as described in Clause 13.

[0026] Clause 15. The statistical likelihood of observing coverage at each genomic locus is the p-value based on the mean and standard deviation defining the distribution of the genomic locus, as described in Clause 9 or any other clause subordinate thereto.

[0027] Clause 16. The contamination metric is the truncated p-value product of the p-values ​​of the cfDNA sample in the genomic locus, as described in Clause 15.

[0028] Clause 17. The contamination threshold is determined in the manner of Clause 9 or any other clause subordinate thereto, in order to achieve a set specificity for the contamination model.

[0029] Clause 18. The method described in Clause 17, wherein the specificity is one of 95%, 96%, 97%, 98%, 99%, and 99.5%.

[0030] Clause 19. A method for detecting leukocyte (WBC) contamination in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: determining coverage for each set of genomic locus feature features as contamination markers, based on the count of sequence reads of cfDNA fragments overlapping with the genomic locus; determining the statistical likelihood of observing coverage at a genomic locus for each set of genomic locus feature features, based on the distribution of coverage for the genomic locus generated from a purified cfDNA sample; generating contamination metrics for the test sample by summing the statistical likelihoods across multiple contamination markers; determining whether the test sample has WBC contamination if the contamination metrics for the test sample cross a contamination threshold; and implementing one or more corrective actions in response to the determination that the test sample has WBC contamination.

[0031] Clause 20. A set of feature quantities for a genomic locus is identified in accordance with the method described in Clause 1 or any other clause subordinate thereto, as described in Clause 19 or any other clause subordinate thereto.

[0032] Clause 21. The distribution of coverage of genomic loci is generated in accordance with the method described in Clause 9 or any other clause subordinate thereto, as described in Clause 19 or any other clause subordinate thereto.

[0033] Clause 22. The contamination threshold is identified in accordance with the method described in Clause 19 or any other clause subordinate thereto.

[0034] Clause 23. A purified cfDNA sample consists essentially of sequence reads corresponding to cfDNA fragments, as described in Clause 19 or any other clause subordinate thereto.

[0035] Clause 24. The statistical likelihood of observing coverage at each genomic locus is the z-score based on the mean and standard deviation defining the distribution of the genomic locus, as described in Clause 19 or any other clause subordinate thereto.

[0036] Clause 25. Contamination metrics are calculated by summing the absolute values ​​of the z-scores, as described in Clause 24 or any other clause subordinate thereto.

[0037] Clause 26. Determining whether the contamination metric for a test specimen crosses the contamination threshold is the method of Clause 24 or any other clause subordinate thereto, including determining whether the contamination metric exceeds the contamination threshold.

[0038] Clause 27. The statistical likelihood of observing coverage at each genomic locus is the p-value based on the mean and standard deviation defining the distribution of the genomic locus, as described in Clause 19 or any other clause dependent thereon.

[0039] Clause 28. Contamination metrics are the truncated p-value product of cfDNA samples in genomic loci, as described in Clause 27 or any other clause subordinate thereto.

[0040] Clause 29. Determining whether the contamination metric for a test sample crosses the contamination threshold is the method of Clause 27 or any other clause subordinate thereto, including determining whether the contamination metric is below the contamination threshold.

[0041] Clause 30. Corrective actions include any combination of the following: notifying the healthcare provider that the test sample has WBC contamination; discarding the test sample; labeling the test sample as contaminated; notifying the healthcare provider to collect subsequent test samples; notifying healthcare providers of potential sources of contamination that may contain WBC; and retaining test samples from downstream analyses which may include cancer classification.

[0042] Clause 31. A method for training a deconvolution model configured to deconvolve the proportion of tissue type contribution in cell-free DNA (cfDNA) samples, the method comprising: obtaining a set of cfDNA samples from different tissue types, each cfDNA sample comprising methylation sequence reads corresponding to a cfDNA fragment; determining, for each cfDNA sample, methylation features at each genomic locus in an initial set of genomic loci based on the methylation sequence reads; generating a feature vector for each sample based on the methylation features across the initial set of genomic loci; and training a deconvolution model to predict the proportion of tissue type based on the feature vectors from the samples.

[0043] Clause 32. The method of Clause 31 or any other clause subordinate thereto, further comprising: identifying a feature set for each genomic locus based on the information gain of methylation features for each genomic locus based on a trained deconvolution model; and retraining the deconvolution model with updated feature vectors based on the feature set for each genomic locus.

[0044] Clause 33. The histological types consist of the groups erythrocyte progenitor type, megakaryocyte type, monocyte type, granulocyte type, lymphocyte type, epithelial cell type, vascular endothelial cell type, hepatocyte type, adipocyte type, endocrine cell type, and myocyte type, as described in Clause 31 or any other clause subordinate thereto.

[0045] Clause 34. The histological types are those comprising the groups of circulating myeloid tissue types, non-circulating myeloid tissue types, lymphocyte cell types, epithelial cell types, vascular endothelial cell types, and hepatocyte types, as described in Clause 31 or any other clause subordinate thereto.

[0046] Clause 35. Each cfDNA sample is purified according to the method described in Clause 31 or any other clause subordinate thereto.

[0047] Clause 36. Each genomic locus covers at least one CpG site, as described in Clause 31 or any other clause subordinate thereto.

[0048] Clause 37. The methylation feature at each genomic locus is one of the following: the methylation density across methylated sequence reads of a cfDNA sample at the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly methylated and overlap with the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly unmethylated and overlap with the genomic locus; and the count or proportion of methylated sequence reads that have a specific methylation variant at the genomic locus, as described in Clause 31 or any other clause subordinate thereto.

[0049] Clause 38. A deconvolution model is a machine learning model as described in Clause 31 or any other clause subordinate thereto.

[0050] Clause 39. A trained deconvolution model is configured to output predictions of tissue type proportions, where each tissue type proportion is the percentage of tissue type contribution to the cfDNA fragments in a cfDNA sample, as described in Clause 31 or any other clause subordinate thereto.

[0051] Clause 40. Information gain refers to the discriminative power of a methylation feature of each genomic locus in distinguishing between two tissue types, as described in Clause 31 or any other clause subordinate thereto.

[0052] Clause 41. A method for training a contamination model to detect the contamination of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: obtaining a set of cfDNA samples from a non-cancer subject, each cfDNA sample containing methylated sequence reads corresponding to a cfDNA fragment; determining, for each cfDNA sample, methylation features at each genomic locus in a set of genomic locus features based on the methylated sequence reads; generating a feature vector for each cfDNA sample based on methylation features across the set of genomic locus features; applying a trained deconvolution model to the feature vectors of each cfDNA sample to predict the proportion of tissue types for each cfDNA sample; and constructing a distribution of the proportion of tissue types for each cfDNA sample, wherein the contamination model includes the distribution.

[0053] Clause 42. A set of feature vectors for a genomic locus is identified and / or a deconvolution model is trained in accordance with the method described in Clause 31 or any other clause subordinate thereto.

[0054] Clause 43. Each cfDNA sample is purified according to the method described in Clause 41 or any other clause subordinate thereto.

[0055] Clause 44. Each genomic locus covers at least one CpG site, as described in Clause 41 or any other clause subordinate thereto.

[0056] Clause 45. The methylation feature at each genomic locus is one of the following: the methylation density across methylated sequence reads of a cfDNA sample at the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly methylated and overlap with the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly unmethylated and overlap with the genomic locus; and the count or proportion of methylated sequence reads that have a specific methylation variant at the genomic locus, as described in Clause 41 or any other clause subordinate thereto.

[0057] Clause 46. The distribution is a multivariate distribution in a multivariate vector space where the number of variables is equal to the number of organizational types, as described in Clause 41 or any other clause subordinate thereto.

[0058] Clause 47. A method for detecting leukocyte (WBC) contamination in a test sample containing methylated sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: determining methylation features at each genomic locus in a set of genomic locus features based on methylated sequence reads; generating a feature vector based on methylation features across the set of genomic locus features; applying a trained deconvolution model to the feature vector of the test sample to predict the proportion of tissue types for the test sample; and generating a contamination metric for the test sample based on the distance of the proportion of tissue types for the test sample to the distribution of the proportion of tissue types generated from cfDNA samples from non-cancer subjects, wherein the contamination metric indicates the likelihood that the cfDNA sample has WBC contamination; determining whether the test sample has WBC contamination if the contamination metric for the test sample crosses a contamination threshold; and implementing one or more corrective actions in response to the determination that the test sample has WBC contamination.

[0059] Clause 48. A set of feature vectors for a genomic locus is identified and / or a deconvolution model is trained in accordance with the method described in Clause 31 or any other clause subordinate thereto, as described in Clause 47 or any other clause subordinate thereto.

[0060] Clause 49. The deconvolution model is trained in accordance with the method described in Clause 41 or any other Clause subordinate thereto, as described in Clause 47 or any other Clause subordinate thereto.

[0061] Clause 50. Each genomic locus covers at least one CpG site, as described in Clause 47 or any other clause subordinate thereto.

[0062] Clause 51. The methylation feature at each genomic locus is one of the following: the methylation density across methylated sequence reads of a cfDNA sample at the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly methylated and overlap with the genomic locus; the count or proportion of methylated sequence reads of a cfDNA sample that are highly unmethylated and overlap with the genomic locus; and the count or proportion of methylated sequence reads that have a specific methylation variant at the genomic locus, as described in Clause 47 or any other clause subordinate thereto.

[0063] Clause 52. Distance is the Mahalanobis distance based on the distribution of proportions of organizational types, as described in Clause 47 or any other clause subordinate thereto.

[0064] Clause 53. Contamination metrics are p-values, as described in Clause 47 or any other clause subordinate thereto.

[0065] Clause 54. Determining whether the contamination metric for a test sample crosses the contamination threshold is the method of Clause 53, which includes determining whether the contamination metric is below the contamination threshold.

[0066] Clause 55. The contamination threshold is determined in accordance with the method of Clause 47 or any other clause subordinate thereto in order to achieve a set specificity for the contamination model.

[0067] Clause 56. Corrective actions include any combination of the following: notifying healthcare providers that a test sample has WBC contamination; discarding the test sample; labeling the test sample as contaminated; notifying healthcare providers to collect subsequent test samples; notifying healthcare providers of potential sources of contamination that may contain WBC; and retaining test samples from downstream analyses which may include cancer classification.

[0068] Clause 57. A method for training a contamination model for detecting leukocyte (WBC) contamination in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising obtaining a first set of cfDNA samples and a second set of WBC samples, each sample containing sequence reads corresponding to DNA fragments, determining the mean coverage of sequence reads overlapping with an initial set of genomic loci for each sample in the first and second sets, and determining the overlapping with genomic loci normalized by the mean coverage of the samples for each sample in the first and second sets. A method comprising: determining normalized coverage for each genomic locus in an initial set of genomic loci based on sequence reads; generating a first distribution of coverage for cfDNA samples and a second distribution of coverage for WBC samples for each genomic locus; identifying highly discriminative genomic loci between cfDNA samples and WBC samples based on the coverage distributions; determining a discriminative score for each genomic locus using a two-sample t-test; and determining a set of genomic locus features as a contamination marker based on the discriminative score.

[0069] Clause 58. Each cfDNA sample consists essentially of sequence reads corresponding to cfDNA fragments, and each WBC sample consists essentially of sequence reads corresponding to DNA fragments excreted from WBCs, as described in Clause 57 or any other clause subordinate thereto.

[0070] Clause 59. Each genomic locus comprises at least one CpG site, as described in Clause 57 or any other clause subordinate thereto.

[0071] Clause 60. The method of Clause 57 or any other clause subordinate thereto, further comprising pruning one or more genomic loci with highly variable coverage across samples in order to generate an initial set of genomic loci.

[0072] Clause 61. The method of Clause 57 or any other clause subordinate thereto, further comprising pruning one or more genomic loci with coverage that is non-normally distributed across samples.

[0073] Clause 62. Identifying highly distinctive genomic loci between cfDNA samples and WBC samples based on coverage distribution, as described in Clause 57 or any other clause subordinate thereto, comprising, for each genomic locus, evaluating the distance between a first coverage distribution for cfDNA samples and a second coverage distribution for WBC samples.

[0074] Clause 63. The distance for each genomic locus is based on the area under the curve (AUC) for the binary classification between coverage-based cfDNA and WBC at the genomic locus, as described in Clause 62.

[0075] Clause 64. Distances for each genomic locus are further based on the Hodges-Lehmann estimator, as described in Clause 63.

[0076] Clause 65.2 A sample t-test is a method of Clause 57 or any other sub-clause that evaluates the discriminability of paired cfDNA and WBC samples for each subject.

[0077] Clause 66. The distinctive score is determined by the method described in Clause 57 or any other clause subordinate thereto, based on the Bonferroni amendment.

[0078] The method of Clause 67 or any other clause subordinate thereto, further comprising pruning any genomic loci having a non-normal distribution of coverage for cfDNA samples or a non-normal distribution of coverage for WBC samples.

[0079] The method of Clause 68, further comprising pruning any genomic loci having a difference between the mean of the coverage distribution for cfDNA samples and the mean of the coverage distribution for WBC samples, which is below a threshold.

[0080] Clause 69. The method of Clause 57 or any other clause subordinate thereto, which includes determining a set of features of a genomic locus as a contaminant marker based on a discriminative score, and identifying genomic loci that have a discriminative score exceeding a threshold.

[0081] Clause 70. Determining a set of genomic locus features as a contaminant marker based on a discriminative score is the method of Clause 57 or any other clause subordinate thereto, which includes ranking an initial set of genomic loci based on a discriminative score and selecting the top number of genomic loci from the ranking to form a set of genomic locus features.

[0082] Clause 71. A method for detecting leukocyte (WBC) contamination in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: determining the mean coverage of sequence reads overlapping with a set of genomic loci; determining normalized coverage for each genomic locus in the set of genomic loci based on sequence reads overlapping with genomic loci normalized by the mean coverage of the sample; applying a contamination model for determining contamination metrics as the contribution rate of WBC efflux DNA to a test sample that maximizes the likelihood of observing normalized coverage across the set of features, based on the distribution of coverage for cfDNA and WBC samples for the set of genomic loci; determining whether a test sample has WBC contamination if the contamination metrics for the test sample cross a contamination threshold; and implementing one or more corrective actions in response to the determination that the test sample has WBC contamination.

[0083] Clause 72. The method of Clause 71 or any other clause thereof, wherein a set of feature vectors for genomic loci is identified, a contamination model is trained, and / or a coverage distribution for cfDNA and WBC samples is generated, in accordance with the method of Clause 57 or any other clause thereof.

[0084] Clause 73. A contamination model is provided in the manner of Clause 71 or any other clause subordinate thereto, including for each genomic locus a first distribution of coverage for cfDNA samples and a second distribution of coverage for WBC samples.

[0085] Clause 74. Determining whether the contamination metric for a test specimen crosses the contamination threshold is the method of Clause 71 or any other clause subordinate thereto, including determining whether the contamination metric is below the contamination threshold.

[0086] Clause 75. The contamination threshold is determined in accordance with the method of Clause 71 or any other clause subordinate thereto, in order to achieve a set specificity for the contamination model.

[0087] Clause 76. Corrective actions include any combination of the following: notifying healthcare providers that a test sample has WBC contamination; discarding the test sample; labeling the test sample as contaminated; notifying healthcare providers to collect subsequent test samples; notifying healthcare providers of potential sources of contamination that may contain WBC; and retaining test samples from downstream analyses which may include cancer classification.

[0088] Clause 77. A method for determining a source of leukocyte (WBC) contamination in one or more samples containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: obtaining a first set of samples from a first sample processing workflow; and obtaining a second set of samples from a second sample processing workflow, wherein the second sample processing workflow includes a first protocol different from the first sample processing workflow, a first clinical product different from the first sample processing workflow, a first sequencing device, or a combination thereof; and applying a contamination model to the first set of samples and the second set of samples to determine contamination metrics for each sample. A method comprising: determining a first aggregate metric for a first sample processing workflow based on a combination of contamination metrics for a first set of samples; determining a second aggregate metric for a second sample processing workflow based on a combination of contamination metrics for a second set of samples; determining that the source of WBC contamination originates from a first protocol, a first clinical product, a first sequencing device, or a combination thereof, based on the determination that the second aggregate metric is greater than the first aggregate metric; and implementing one or more corrective actions to mitigate WBC contamination from the second sample processing workflow.

[0089] Clause 78. The method of Clause 77, wherein the imputed model is trained in accordance with the method of Clause 9 or any other Clause subordinate thereto, Clause 41 or any other Clause subordinate thereto, or Clause 57 or any other Clause subordinate thereto.

[0090] Clause 79. Corrective actions include any combination of changing the first protocol in a second sample processing workflow to another protocol, changing the first clinical product in a second sample processing workflow to another clinical product, changing the first sequencing device in a second sample processing workflow to another sequencing device, obtaining a third set of samples instead of a second set of sequenced samples from the first sample processing workflow, and discarding the second set of samples obtained from the second sample processing workflow, as described in Clause 77 or any other clause subordinate thereto.

[0091] Clause 80. The method of Clause 77 or any other clause subordinate thereto, wherein the first aggregated contamination metric is representative of the contamination metrics of a first set of samples, and the second aggregated contamination metric is representative of the contamination metrics of a second set of samples.

[0092] Clause 81. The method as described in Clause 77 or any other clause subordinate thereto, wherein the first aggregated contamination metric is a first proportion or first count of samples in a first set of samples having a corresponding contamination metric exceeding the contamination threshold, and the second aggregated contamination metric is a second proportion or second count of samples in a second set of samples having a corresponding contamination metric exceeding the contamination threshold.

[0093] Clause 82. A method for training a cancer classifier using samples containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, the method comprising: obtaining a first set of samples obtained from a first cohort of healthy subjects and a second set of samples obtained from a second cohort of subjects diagnosed with cancer; applying a contamination model to the first set of samples and the second set of samples to determine contamination metrics for each sample indicating the amount of leukocyte (WBC) contamination in the samples; determining one or more contaminated samples having corresponding contamination metrics exceeding a contamination threshold; filtering the contaminated samples, the filtering may include discarding the contaminated samples; determining feature vectors for each remaining sample based on the sequence reads of the samples; and training a cancer classifier with the feature vectors for the remaining samples, the trained cancer classifier being configured to predict the likelihood of cancer presence based on input feature vectors derived based on sequence reads in test cfDNA samples.

[0094] Clause 83. The impregnated model is trained in accordance with the method described in Clause 9 or any other clause subordinate thereto, Clause 41 or any other clause subordinate thereto, or Clause 57 or any other clause subordinate thereto, as described in Clause 82.

[0095] Clause 84. A cancer classifier is a machine learning model, as described in Clause 82 or any other clause subordinate thereto.

[0096] Clause 85. A cancer classifier is trained to predict a binary prediction between the presence or absence of cancer, as described in Clause 82 or any other clause subordinate thereto.

[0097] Clause 86. A second set of samples comprises one or more subsets of the samples, each subset of the samples obtained from a subject diagnosed with one of several types of cancer, as described in Clause 82 or any other sub-clauses thereof.

[0098] Clause 87. The method of Clause 86, wherein the cancer classifier is trained to predict multi-class predictions as the likelihood of the presence of multiple cancer types.

[0099] Clause 88. For each sample, the sequence reads are methylated sequence reads, and the feature vectors are based on the methylated sequence reads having a useful methylation pattern, as described in Clause 82 or any other clause subordinate thereto.

[0100] Clause 89. A non-temporary computer-readable storage medium that stores instructions causing one or more computer processors to perform any of the methods described in Clauses 1 to 88 when executed by one or more computer processors.

[0101] Clause 90. A system comprising one or more computer processors and a non-temporary computer-readable storage medium as described in Clause 89.

[0102] Clause 91. A therapeutic kit comprising: a collection container for collecting a DNA sample from a subject; one or more reagents for optionally isolating DNA fragments in the DNA sample; one or more contamination probes for optionally targeting one or more genomic loci selected from Table 1; and a non-temporary computer-readable storage medium as defined in Clause 89. [Brief explanation of the drawing]

[0103] [Figure 1] This is an illustrative flowchart illustrating the overall workflow for cancer classification of a sample according to one or more embodiments. [Figure 2A] This is an exemplary flowchart illustrating a process for sequencing cell-free (cf) DNA fragments to obtain a methylation state vector, according to one or more embodiments. [Figure 2B] Figure 2A illustrates the process of sequencing cell-free (cf) DNA fragments to obtain a methylation state vector according to one or more embodiments. [Figure 3A]A flowchart illustrating the identification of features for coverage-based WBC contamination detection according to one or more embodiments is shown. [Figure 3B] A flowchart is provided illustrating the training of a contamination model for predicting coverage-based WBC contamination across feature sets of genomic loci using one or more embodiments. [Figure 3C] A flowchart illustrating the prediction of WBC contamination using a contamination model in one or more embodiments is shown. [Figure 4A] A flowchart illustrating the training of a deconvolution model for methylation-based WBC contamination detection in one or more embodiments is shown. [Figure 4B] A flowchart illustrating the training of a contamination model for predicting WBC contamination based on the proportion of predicted tissue types, using one or more embodiments, is shown. [Figure 4C] A flowchart illustrating the prediction of WBC contamination using a contamination model in one or more embodiments is shown. [Figure 5A] A flowchart illustrating the feature identification and training of a contamination model for quantitative coverage-based WBC contamination detection, according to one or more embodiments, is shown. [Figure 5B] A flowchart illustrating the prediction of WBC contamination using a contamination model in one or more embodiments is shown. [Figure 6A] This is an exemplary flowchart illustrating a process for generating a control group data structure for determining highly informative methylated fragments, according to one or more embodiments. [Figure 6B] This is an exemplary flowchart illustrating a process for determining highly informative methylated fragments based on a control group data structure, according to one or more embodiments. [Figure 7A] This is an illustrative flowchart illustrating the process of training a cancer classifier according to one or more embodiments. [Figure 7B] This document illustrates the generation of feature vectors used to train a cancer classifier, according to one or more embodiments. [Figure 8A] An exemplary flowchart of an apparatus for sequencing nucleic acid samples according to one or more embodiments is shown. [Figure 8B] This is an exemplary block diagram of an analysis system according to one or more embodiments. [Figure 9A] A volcano plot of coverage ratios for an initial set of genomic loci under consideration is shown, according to an exemplary embodiment. [Figure 9B] This illustrates a visual comparison of differential coverage at selected genomic loci between cfDNA samples and WBC samples using an exemplary embodiment. [Figure 10] An exemplary embodiment shows a single genomic locus in mitochondrial DNA that exhibits a significant difference in representation between a cfDNA sample and a WBC sample. [Figure 11] This example demonstrates a comparison of significant genomic loci in a set of genomic locus features compared to a randomly selected candidate set. [Figure 12] An exemplary embodiment shows a pair of WBC and cfDNA samples from the same subject, evaluated across a set of genomic locus features. [Figure 13] This shows a sample identified as containing contaminated WBCs by coverage-based WBC contamination detection according to an exemplary embodiment. [Figure 14] This shows a sample (top 1%) having the longest fragment identified as contaminated WBC by coverage-based WBC contamination detection, according to an exemplary embodiment. [Figure 15] The contribution rates of various tissue types are shown for three different sample types according to exemplary embodiments: cfDNA samples, purified plasma samples, and serum samples containing only WBC DNA. [Figure 16] This illustrates the information gain across a set of genomic loci evaluated for inclusion as part of a deconvolution model, using an exemplary embodiment. [Figure 17]A box plot illustrating the predictive power of a methylation-based WBC contamination detection methodology according to an exemplary embodiment is shown. [Figure 18] An exemplary embodiment shows a titration sample mixed with cfDNA and WBC serum. [Figure 19] A titration sample having a Mahalanobis distance calculated according to a second methodology, as in an exemplary embodiment, is shown. [Figure 20] An exemplary embodiment shows a titration sample for which WBC contamination has been quantified. [Modes for carrying out the invention]

[0104] The drawings depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following considerations that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein.

[0105] I. Overview Early detection and classification of cancer are crucial skills. Detecting cancer before it becomes symptomatic is beneficial for all involved parties, including patients, physicians, and those close to the patient. For patients, early cancer detection increases the likelihood of a favorable outcome; for physicians, it opens up more treatment pathways that can lead to favorable outcomes; and for those close to the patient, early cancer detection increases the chances of not losing friends and family due to the disease.

[0106] In recent years, early cancer detection technologies have advanced in the direction of analyzing gene fragments (e.g., DNA) in human blood to determine whether any of these gene fragments originate from cancer cells. These novel technologies allow physicians to identify the presence of cancer in patients that might otherwise be undetectable in conventional screening processes. For example, consider a person at high risk of breast cancer. Traditionally, this person visits a doctor regularly to have a mammogram (e.g., an X-ray image) taken, which the doctor uses to identify cancerous tissue. Unfortunately, even with a high-resolution mammogram, once the tumor is about 1 millimeter in size, the doctor can only identify the tumor. This means that the cancer has been present in the person for some time, undiagnosed and untreated. Such visual diagnosis is typical for most cancers; that is, they can only be identified after they have grown to a sufficient size and become identifiable with certain imaging techniques.

[0107] Cancer detection using analysis of gene fragments in a patient's blood, for example, mitigates this problem. For instance, cancer cells begin shedding DNA fragments into the human bloodstream as soon as they are formed. This occurs when there are very few cancer cells and before they are visible with imaging techniques. Therefore, in a suitable manner, a system that analyzes DNA fragments in the bloodstream can identify the presence of cancer in a person based on the shedding cancer DNA fragments, and more importantly, this system can identify the presence of cancer before it is identified using more traditional cancer detection techniques.

[0108] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing (NGS) technology. NGS is a broad group of technologies that enable high-throughput sequencing of genetic material. As will be discussed in more detail herein, NGS mainly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the laboratory method required to prepare DNA fragments for sequencing, sequencing is the process of reading ordered nucleotides in a sample, and data analysis is the processing and analysis of genetic information in the sequence data to identify the presence of cancer.

[0109] While these steps in NGS can help enable early cancer detection, they also introduce their own complex and detrimental problems to cancer detection. Therefore, improvements to sample preparation, DNA sequencing, and / or data analysis, including preprocessing, algorithmic processing, and summarizing or presenting predictions or conclusions, more generally lead to improvements in cancer detection techniques and early cancer detection.

[0110] For example, (1) problems arising in sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Improving these issues would result in better genetic data for cancer detection. Similarly, (2) problems arising in sequencing include, for example, errors in accurate transcription of fragments (e.g., reading "A" instead of "C"), inaccurate or difficult fragment assembly and duplication, varying coverage uniformity, sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Again, improving any of these issues would result in better genetic data for cancer detection.

[0111] (3) The problems in data analysis are the most difficult and complex. The challenges stem from the enormous amount of data created by NGS sequencing technology. The resulting gene datasets are typically in the terabyte range, and effectively and efficiently analyzing this volume of data is required both procedurally and computationally. For example, analyzing NGS sequencing involves several baseline processing steps, such as aligning reads with each other, aligning and mapping reads to a reference genome, identifying and calling mutant genes, identifying and calling abnormally methylated genes, and generating functional annotations. Performing any of these processes on terabytes of gene data is computationally expensive even with the most powerful computer architectures, and is completely impossible for ordinary human thinking. In addition, gene sequencing data obtained from error-prone sample preparation and sequence reading processes may result in a large portion of the resulting gene data being of low quality or unsuitable for cancer identification. For example, large amounts of gene data may contain contaminated samples, transcription errors, mismatched regions, and overexpressed regions, which may make them unsuitable for high-precision cancer detection. Identifying and explaining low-quality genetic data across the vast amounts of genetic data obtained from NGS sequencing is procedurally and computationally rigorous to achieve, and practically impossible for human thought. Overall, any process created that results in more efficient processing of large array sequencing data would be an improvement in cancer detection using NGS sequencing.

[0112] Finally, and perhaps most importantly, accurately identifying highly informative DNA from NGS data to identify the presence of cancer is also difficult (even more so in the context of early cancer detection). To be effective, algorithms must compensate for errors generated by, for example, sample preparation and sequencing, and overcome the problems of large-scale data analysis associated with NGS technology. In other words, designing one or more machine learning models or other computational algorithms that enable early cancer detection based on next-generation sequencing technology must be configured to take into account the problems that these technologies present. Some of these technologies and models are discussed below, and specific improvements to state-of-the-art technologies and models are discussed further.

[0113] One specific challenge arises during the sample preparation stage due to WBC efflux DNA. Cancer signals are typically confused with WBC efflux signals, which can interfere with cancer detection when analyzed in conjunction with cfDNA in plasma. This signal confusion can lead to false positive calls for cancer. Computer models implemented in WBC contamination detection aim to assess whether a sample is contaminated with WBC efflux DNA and can further determine the level of contamination. If a WBC-contaminated sample is detected, the analytical system can take corrective action. For example, the analytical system can eliminate the use of such samples in training a cancer classification model, thereby improving the accuracy and precision of the cancer classification model. To extend further, the analytical system can perform WBC contamination detection from an initial set of training samples to determine any WBC-contaminated samples to remove from the set of training samples. The filtered set of training samples (without WBC contamination) can then be used to train the model. As another example, the analytical system can refrain from using contaminated test samples in cancer classification, thereby reducing the false positive rate of cancer predictions. If a contaminated sample is detected, further corrective action may be taken to address the contamination. Corrective actions may include identifying the source of contamination by testing various conditions in the sample preparation process to determine the conditions that contribute to WBC contamination. Removing the source of contamination improves the physical assay process and sample handling. Once the source of contamination is identified, measures are taken to remove the source or minimize contamination in the sample handling workflow. Another corrective action may be to physically discard the sample.

[0114] Training of the machine learning models described herein (including imputation models, cancer classifiers, any other neural networks, and any other models referenced herein) involves performing one or more non-mathematical operations, or at least partially implementing non-mathematical functions by a machine or computing system, examples of which include, but are not limited to, data loading operations, data storage operations, data toggle or modification operations, non-temporary computer-readable storage medium modification operations, metadata removal or data cleansing operations, data compression operations, protein structure modification operations, image modification operations, noise application operations, and noise reduction operations. Thus, training of the machine learning models described herein may be based on or include mathematical concepts, and is not simply limited to performing mathematical calculations, mathematical operations, or acts of calculating variables or numbers using mathematical methods.

[0115] Similarly, it should be noted that training these models described herein is practically impossible to perform in human thought. The models are inherently complex and involve a vast number of weights and parameters linked through one or more complex functions. Training and / or deploying such models involves so many operations that it is impossible to perform them solely in human thought, nor even with the help of pen and paper. In such embodiments, the operations can number in the hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions. In addition, the training data has hundreds, thousands, tens of thousands, hundreds of thousands, millions, or billions of sequence reads, and each sequence read may further contain anywhere from hundreds to thousands of nucleotides. Thus, such models are inevitably rooted in computer technology for their implementation and use.

[0116] IA cancer classification workflow Figure 1 is an exemplary flowchart illustrating an overall workflow 100 for cancer classification of a sample, according to one or more embodiments. The workflow 100 is carried out by one or more entities, including, for example, healthcare providers, sequencing equipment, and analytical systems. The purpose of the workflow includes cancer detection and / or monitoring in an individual. From a healthcare perspective, workflow 100 can serve to complement other existing cancer diagnostic tools. Workflow 100 may function to provide early cancer detection and / or regular cancer monitoring to better inform treatment plans for individuals diagnosed with cancer. The overall workflow 100 may include additional steps / fewer steps than those shown in Figure 1.

[0117] A healthcare provider performs sample collection 110. An individual undergoing cancer classification visits the healthcare provider. The healthcare provider collects a sample to perform cancer classification. Examples of biological samples include, but are not limited to, a tissue biopsy of the subject, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. In one or more embodiments, sample collection is minimally invasive or non-invasive. The sample contains genetic material belonging to the individual and can be extracted and sequenced for cancer classification. Once the sample is collected, it is provided to a sequencing device. Along with the sample, the healthcare provider may collect other information related to the individual, such as biological sex, age, race, smoking status, other health indicators, and any previous diagnoses.

[0118] The sequencing instrument performs sample sequencing 120. A laboratory clinician may perform one or more processing steps on the sample during sequencing preparation. Once preparation is complete, the clinician loads the sample into the sequencing instrument. Examples of instruments used for sequencing are further illustrated in conjunction with Figures 8A and 8B. Sequencing instruments generally extract and isolate the nucleic acid fragments to be sequenced in order to determine the sequence of nucleic acid bases corresponding to the fragments. Sequencing may also involve amplification of nucleic acid material. Different sequencing processes include Sanger sequencing, fragment analysis, and other next-generation sequencing techniques. Next-generation sequencing can yield high-throughput sequence data, e.g., 10,000, 100,000, 1,000,000, or 10,000,000 sequence reads, each sequence read may be of a length such as 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, or 300 bp. Therefore, high-throughput sequencing data are of a size that is impractical for analysis by human thought. Sequencing can be whole-genome sequencing or targeted sequencing by a targeted panel. In relation to DNA methylation, bisulfite sequencing (e.g., further described in Figures 2A and 2B) can determine the methylation state via the bisulfite conversion of unmethylated cytosine at CpG sites. Sample sequencing 120 produces sequences of multiple nucleic acid fragments in the sample. In one or more embodiments, the sequences may include methylation state vectors, each methylation state vector describing the methylation state of CpG sites on the fragment.

[0119] The analysis system performs pre-analysis processing 130. An exemplary analysis system is illustrated in Figure 8B. Pre-analysis processing 130 may include, but is not limited to, deduplication of sequence reads, determination of coverage-related metrics, determination of contamination in the sample (including determination of WBC contamination), removal of contaminated fragments, and calling for sequencing errors.

[0120] The analysis system performs one or more analyses 140. The analysis is the application of one or more trained statistical analyses or models to predict at least the cancerous status of the individual from which the sample originates. Various genetic features such as CpG site methylation, single nucleotide polymorphism (SNP), insertion or deletion (indel), and other types of gene mutations may be evaluated and considered. In relation to methylation, analysis 140 may include contamination detection 142 (e.g., further described in Figures 3A-3C, 4A-4C, and 5A and 5B), highly informative methylation identification 144 (e.g., further described in Figures 6A and 6B), feature extraction 146 (e.g., further described in Figures 7A and 7B), and applying a cancer classifier 148 to determine the cancer prediction (e.g., Figures 7A and 7B). The cancer classifier 148 determines the cancer prediction by inputting the extracted features. The cancer prediction may be a label or a value. The labels can indicate specific cancer conditions; for example, a binary label can indicate the presence or absence of cancer, and a multiclass label can indicate one or more cancer types from multiple cancer types being screened. The values ​​may indicate the likelihood of a specific cancer condition, e.g., the likelihood of cancer, and / or the likelihood of a specific cancer type.

[0121] The analysis system returns a prediction 150 to the healthcare provider. The healthcare provider may establish or adjust the treatment plan based on the cancer prediction. Treatment optimization is further described in Section IV.C. Treatment. In some embodiments, the analysis system may leverage a cancer classification workflow for prognosis determination, treatment individualization, treatment evaluation, and cancer status monitoring.

[0122] Overview of IB Methylation In accordance with this specification, individual-derived cfDNA fragments are processed and sequenced, for example by converting unmethylated cytosine to uracil, and the sequences are read by comparing them to a reference genome to identify the methylation status at specific CpG sites within the DNA fragment. Each CpG site may be methylated or unmethylated. By identifying highly informative methylated fragments compared to healthy individuals, insights into the subject's cancer status may be provided. As is well known in the art, abnormal DNA methylation (compared to healthy controls) can cause a variety of effects, which may contribute to cancer. Identifying highly informative methylated cfDNA fragments presents several challenges. Firstly, determining highly informative methylated DNA fragments can be weighted compared to a group of control individuals, and consequently, if the number of control individuals is small, the determination loses reliability due to statistical variability within the smaller size of the control group. In addition, methylation status may vary among groups of control individuals, which can be difficult to explain when determining the DNA fragments to be methylated that are highly informative. In another respect, cytosine methylation at a CpG site can have a causal effect on subsequent methylation at other CpG sites. Encapsulating this dependency can itself become another challenge.

[0123] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation can occur in cytosine and guanine dinucleotides, which are referred to herein as “CpG sites.” In other cases, methylation can occur in cytosine or other nucleotides that are not part of a CpG site, but these are less likely to occur. For clarity, in this disclosure, methylation is discussed with reference to CpG sites. Highly informative DNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate cancerous conditions. Throughout this disclosure, hypermethylation and hypomethylation can be characterized in DNA fragments if the DNA fragment contains more than a threshold number of CpG sites, and more than that threshold percentage of CpG sites are methylated or demethylated.

[0124] The principles described herein may be equally applicable to the detection of methylation related to non-CpG, including non-cytosine methylation. In such embodiments, the wet laboratory assay used to detect methylation may differ from that described herein. Furthermore, the methylation state vectors discussed herein may generally contain elements that are sites where methylation has occurred or not (even if these sites are not clearly CpG sites). With such substitution, the remainder of the process described herein may be the same, and as a result, the inventive concepts described herein can be applied to those other forms of methylation.

[0125] IC definition The term "cell-free nucleic acid" or "cfNA" refers to nucleic acid fragments that circulate within an individual's body (e.g., blood) and originate from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells). The term "cell-free DNA" or "cfDNA" refers to deoxyribonucleic acid fragments that circulate within an individual's body (e.g., blood). In addition, cfNA or cfDNA within an individual's body may originate from other non-human sources.

[0126] The terms “genomic nucleic acid,” “genomic DNA,” or “gDNA” refer to nucleic acid molecules or deoxyribonucleic acid molecules obtained from one or more cells. In various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or tumor cells (e.g., biopsy samples). In some embodiments, gDNA can be extracted from cells derived from blood cell lineages such as leukocytes.

[0127] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments originating from tumor cells or other types of cancer cells that may be released into the bodily fluids of an individual (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dead cells, or that may be actively released by living tumor cells.

[0128] The terms "DNA fragment," "fragment," or "DNA molecule" generally refer to any deoxyribonucleic acid fragment, i.e., cfDNA, gDNA, ctDNA, etc.

[0129] The terms "highly informative fragment," "highly informative methylated fragment," or "fragment with a highly informative methylation pattern" refer to fragments with abnormal methylation of CpG sites. Abnormal methylation of fragments can be determined using probabilistic models, and the unexpectedness of observing the methylation pattern of fragments in a control group can be identified.

[0130] The terms “UFXM” or “UfxM” refer to hypomethylated or hypermethylated fragments. Hypomethylated and hypermethylated fragments refer to fragments having methylation (e.g., 90%) or unmethylated CpG sites (e.g., 5) that exceed a certain threshold percentage.

[0131] The term "information score" refers to a score for a CpG site based on the overlap between a large number of highly informative fragments (or, in some embodiments, UFXMs) from a sample and its CpG site. The information score is used in relation to the characterization of the sample for classification.

[0132] As used herein, the terms “about” or “approximately” may mean within an acceptable margin of error for a particular value as determined by those skilled in the art, which may depend in part on how the value is measured or determined, for example, on the limits of the measuring system. For example, “about” may mean within one standard deviation or two or more standard deviations, according to the practice of the art. “About” may mean within ±20%, ±10%, ±5%, or ±1% of a given value. The terms “about” or “approximately” may mean within one decimal place, five times, or two times the value. Where a particular value is described in the application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error for that particular value. The term “about” may have a meaning that is generally understood by those skilled in the art. The term “about” may mean ±10%. The term “about” may mean ±5%.

[0133] As used herein, the terms “biological sample,” “patient sample,” or “sample” refer to any sample taken from a subject, which may reflect the biological state associated with the subject and may include cell-free DNA. A sample may be a liquid sample or a solid sample (e.g., a cell or tissue sample). A biological sample may be a bodily fluid such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., from the testes), vaginal flushing fluid, pleural fluid, pericardial fluid, peritoneal fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, discharge from the nipple, or aspirates from different parts of the body (e.g., thyroid, breast). A biological sample may be a fecal sample. A biological sample may include any tissue or substance derived from a living or deceased subject. A biological sample may be a cell-free sample. A biological sample may include nucleic acids (e.g., DNA or RNA) or fragments thereof.

[0134] The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. Nucleic acids in a sample may be cell-free nucleic acids. In various embodiments, the majority of DNA in a biological sample enriched with cell-free DNA (e.g., a plasma sample obtained by a centrifugation protocol) may be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free). Biological samples may be treated to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis), thereby releasing intracellular components into a solution that may further contain enzymes, buffers, salts, surfactants, etc., which can be used to prepare a sample for analysis.

[0135] As used herein, the terms “control,” “control sample,” “reference,” “reference sample,” “normal,” and “normal sample” describe a sample derived from an object that does not have a particular condition or is otherwise healthy. For example, the method disclosed herein can be performed on an object with a tumor, and the reference sample is a sample taken from healthy tissue of the object. The reference sample can be obtained from the object or from a database. The reference may be, for example, a reference genome used to map nucleic acid fragment sequences obtained from sequencing of a sample derived from the object. The reference genome may refer to a haploid or diploid genome from which nucleic acid fragment sequences from biological samples and constituent samples can be aligned and compared. An example of a constituent sample may be DNA from leukocytes obtained from the object. In the case of a haploid genome, there may be only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified, and each heterozygous locus may have two alleles, either of which may allow for matching for alignment to that locus.

[0136] As used herein, the terms “cancer” or “tumor” refer to an abnormal mass of tissue whose growth exceeds and does not coordinate with the growth of normal tissue.

[0137] As used herein, the term “healthy” refers to an object in good health. A healthy object may demonstrate the absence of any malignant or non-malignant disease. A “healthy individual” may have other diseases or conditions unrelated to the condition being assayed that would not normally be considered “healthy.”

[0138] As used herein, the term “methylation” refers to the modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation tends to occur at cytosine and guanine dinucleotides, which are referred to herein as “CpG sites.” In other cases, methylation may occur at cytosine or other nucleotides that are not part of a CpG site, but these are less likely to occur. Highly informative cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate cancerous conditions. Abnormal DNA methylation (compared to a healthy control) can cause a variety of effects, which may contribute to cancer. The principles described herein are equally applicable to the detection of CpG-related and non-CpG-related methylation, including non-cytosine methylation. Furthermore, the methylation state vector may generally contain elements that are vectors of sites where methylation has occurred or not (even if these sites are not clearly CpG sites).

[0139] Where used interchangeably herein, the terms “methylated fragment” or “nucleic acid methylated fragment” refer to a sequence of methylation states of each CpG site in a plurality of CpG sites, determined by methylation sequencing of nucleic acids (e.g., nucleic acid molecules and / or nucleic acid fragments). In a methylated fragment, the location and methylation state of each CpG site in the nucleic acid fragment are determined based on the alignment of sequence reads (e.g., obtained from sequencing of nucleic acids) to a reference genome. A nucleic acid methylated fragment includes the methylation states of each CpG site in a plurality of CpG sites (e.g., a methylation state vector), which identifies the location of the nucleic acid fragment in the reference genome (e.g., identified by the location of a first CpG site in the nucleic acid fragment using a CpG index or another similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of sequence reads to a reference genome based on methylation sequencing of nucleic acid molecules can be performed using a CpG index. As used herein, the term “CpG index” refers to a list of each CpG site in a reference genome, such as the human reference genome, which may be in electronic form (e.g., CpG1, CpG2, CpG3, etc.). The CpG index further includes, for each CpG site in the CpG index, the corresponding genomic location in the corresponding reference genome. Thus, each CpG site in each nucleic acid methylation fragment is indexed to a specific location in its respective reference genome and can be determined using the CpG index.

[0140] As used herein, the term “true positive” (TP) refers to an object having a disease. A “true positive” may refer to an object having a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease. A “true positive” may refer to an object having a disease and being identified as having a disease by the assay or method of this disclosure. As used herein, the term “true negative” (TN) refers to an object that does not have a disease or does not have a detectable disease. A true negative may refer to an object that does not have a disease or detectable disease such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, a non-malignant disease, or a otherwise healthy object. A true negative may refer to an object that does not have a disease or does not have a detectable disease, or is identified as not having such a disease by the assay or method of this disclosure.

[0141] As used herein, the term “reference genome” refers to any specific known, sequenced, or characterized genome, whether partial or complete, of any organism or virus that can be used to reference sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). “Genome” refers to the complete genetic information of an organism or virus as represented by nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of a species’ set of genes. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0142] As used herein, the terms “sequence read” or “read” refer to a nucleotide sequence produced by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (“single-ended read”), or in some cases from both ends of the nucleic acid (e.g., paired-ended read, double-ended read). In some embodiments, sequence reads (e.g., single-ended or paired-ended reads) can be generated from one or both strands of a targeted nucleic acid fragment. The length of a sequence read is often related to the particular sequencing technique. High-throughput methods, for example, provide sequence reads whose size can vary from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads have a typical length, median length, and average length of approximately 15 bp to approximately 900 bp in length (e.g., approximately 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 200 bp, 450 bp, 300 bp, 350 bp, 400 bp, 450 bp, or approximately 500 bp). In some embodiments, sequence reads have a representative, median, or average length of approximately 1,000 bp, 2,000 bp, 5,000 bp, 10,000 bp, or 50,000 bp or more. Nanopore sequencing can provide sequence reads whose size can vary, for example, by tens, hundreds, or thousands of base pairs. Illumina parallel sequencing can provide less variable sequence reads, for example, the majority of which may be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read may correspond to a string of nucleotides from a portion of a nucleic acid fragment (e.g., about 20 to about 150), to a string of nucleotides at one or both ends of a nucleic acid fragment, or to nucleotides from the entire nucleic acid fragment.Sequence reads can be obtained in various ways, for example, by sequencing techniques or by using probes, for example, by hybridization arrays or capture probes, or by amplification techniques, for example, polymerase chain reaction (PCR) or linear amplification using single primers or isothermal amplification.

[0143] As used herein, the term "sequencing" generally refers to any biochemical process that can be used to determine the order of biomacromolecules such as nucleic acids or proteins. For example, sequence data may include all or some of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.

[0144] As used herein, the term “sequencing depth” is used interchangeably with the term “coverage” and refers to the number of times a locus is covered by consensus sequence reads corresponding to specific nucleic acid target molecules aligned to that locus, for example, sequencing depth is equal to the number of specific nucleic acid target molecules covering that locus. This locus may be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as “Y times,” for example, 50 times, 100 times, etc., where “Y” refers to the number of times the locus is covered by sequences corresponding to nucleic acid targets, for example, the number of times independent sequence information covering a particular locus can be obtained. In some embodiments, sequencing depth corresponds to the number of genomes sequenced. Sequencing depth can also be applied to multiple loci or the whole genome, in which case Y may refer to the average number of times the locus, haploid genome, or whole genome is sequenced, respectively. When the average depth is given, the actual depths of different loci included in the dataset can range from one value to another. Ultra-deep sequencing may refer to a sequencing depth of at least 100 times at a locus.

[0145] As used herein, the terms “sensitivity” or “true positive rate” (TPR) refer to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can be characterized by the ability of an assay or method to accurately identify a certain proportion of a population that truly has a disease. For example, sensitivity can be characterized by the ability of a method to accurately identify the number of subjects in a population that has cancer. In another example, sensitivity can be characterized by the ability of a method to accurately identify one or more markers that indicate cancer.

[0146] As used herein, the terms “specificity” or “true negative rate” (TNR) refer to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to accurately identify a certain proportion of a population that does not truly have a disease. For example, specificity can characterize the ability of a method to accurately identify the number of subjects in a population that does not have cancer. In another example, specificity can characterize the ability of a method to accurately identify one or more markers indicating cancer.

[0147] As used herein, the term “Subject” means any living or non-living organism, including but not limited to humans (e.g., males, females, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle (e.g., cows), horses (e.g., horses), goats and sheep (e.g., sheep, goats), pigs (e.g., pigs), camels (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursids (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks, may function as a subject. In some embodiments, the subject is a male or female at any stage (e.g., a male, female, or child). The subjects from whom samples are taken, or who are treated by any of the methods or compositions described herein, may be of any age and may be adults, infants, or children.

[0148] As used herein, the term “tissue” may refer to a group of cells that function together as a single unit. Multiple types of cells may be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but may also refer to tissues derived from different organisms (maternal vs. fetal) or healthy cells vs. tumor cells. The term “tissue” may generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharynx tissue, oropharyngeal tissue). In some embodiments, the terms “tissue” or “type of tissue” may be used to refer to the tissue from which cell-free nucleic acids originate. In one example, viral nucleic acid fragments may originate from blood tissue. In another example, viral nucleic acid fragments may originate from tumor tissue.

[0149] As used herein, the term “genomic” refers to the characteristics of the genome of an organism. Examples of genomic characteristics include, but are not limited to, those relating to the primary nucleic acid sequences of all or part of the genome (e.g., presence or absence of nucleotide polymorphisms, indels, sequence rearrangements, mutation frequencies, etc.), the copy number of one or more specific nucleotide sequences in the genome (e.g., copy number, allele frequency ratio, monochromosome or whole genome ploidy, etc.), the epigenetic status of all or part of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modification, nucleosome arrangement, etc.), and the expression profile of the genome of an organism (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).

[0150] The technical terms used herein are intended to illustrate specific cases and are not intended to limit them. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context specifically indicates otherwise. Furthermore, wherever the terms “including,” “includes,” “having,” “has,” and “with,” or their variations thereof, are used in either the detailed description or / or the claims, such terms are intended to be comprehensive in the same manner as the term “comprising.”

[0151] ID Exemplary Analysis System Figure 8A shows an exemplary flowchart of an apparatus for sequencing nucleic acid samples according to one or more embodiments. This exemplary flowchart includes apparatus such as a sequencer 820 and an analysis system 800. The sequencer 820 and the analysis system 800 may operate in parallel to perform one or more steps in the process.

[0152] In various embodiments, the sequencer 820 receives a concentrated nucleic acid sample 810. As shown in Figure 8A, the sequencer 820 may include a graphical user interface 825 that enables user interaction with specific tasks (e.g., starting or ending sequencing), and one or more loading stations 830 for loading a sequencing cartridge containing the concentrated fragment sample and / or loading the buffers necessary to perform the sequencing assay. Thus, once the user of the sequencer 820 provides the necessary reagents and sequencing cartridge to the loading station 830 of the sequencer 820, the user can start sequencing by interacting with the graphical user interface 825 of the sequencer 820. Once started, the sequencer 820 performs sequencing and outputs sequence reads of the concentrated fragment from the nucleic acid sample 810.

[0153] In some embodiments, the sequencer 820 is communicatively coupled to an analysis system 800. The analysis system 800 includes several computing devices used to process sequence reads for various applications such as evaluation of methylation status at one or more CpG sites, variant calling, or quality control. The sequencer 820 may provide the analysis system 800 with sequence reads in BAM file format. The analysis system 800 can be communicatively coupled to the sequencer 820 via wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analysis system 800 comprises a processor and a non-temporary computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process sequence reads or to perform one or more steps of any of the methods or processes disclosed herein.

[0154] In some embodiments, sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information. The alignment position can generally describe the start and end positions of regions in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. Corresponding to methylation sequencing, the alignment position information can be generalized to indicate the first and last CpG sites included in the sequence read according to its alignment to the reference genome. The alignment position information may further indicate the methylation status and location of all CpG sites in a given sequence read. Regions in the reference genome may be associated with genes or segments of genes. Therefore, the analysis system 800 can label a sequence read with one or more genes that align to that sequence read. In one embodiment, the fragment length (or size) is determined from the start and end positions.

[0155] In various embodiments, for example, when using a paired-end sequencing process, a sequence read consists of a read pair designated R_1 and R_2. For example, the first read R_1 may be sequenced from the first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from the second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be aligned (e.g., in opposite directions) to match the nucleotide bases of a reference genome. The alignment position information derived from the read pair R_1 and R_2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R_1) and an end position in the reference genome corresponding to the end of the second read (e.g., R_2). In other words, the start and end positions in the reference genome can represent possible locations of nucleic acid fragments within the corresponding reference genome. Output files in SAM (Sequence Alignment Map) or BAM (Binary) format are generated and can be output for further analysis.

[0156] Referring now to Figure 8B, which is a block diagram of an analysis system 800 for processing a DNA sample according to one embodiment. The analysis system implements one or more computing devices used for the analysis of the DNA sample. The analysis system 800 includes a sequence processor 840, a sequence database 845, a model database 855, a model 850, a parameter database 865, and a score engine 860. In some embodiments, the analysis system 800 performs some or all of the processes described throughout this disclosure.

[0157] The sequencing processor 840 generates methylation state vectors for the fragments from the sample. For each CpG site on the fragment, the sequencing processor 840 generates a methylation state vector for each fragment that identifies the fragment's location in the reference genome, the number of CpG sites on the fragment (whether methylated, unmethylated, or undetermined), and the methylation state of each CpG site on the fragment, via process 200 in Figure 2A. The sequencing processor 840 can store the methylation state vectors of the fragments in the sequencing database 845. The data in the sequencing database 845 can be organized so that the methylation state vectors from the sample are correlated with each other.

[0158] Furthermore, multiple different models 850 may be stored in the model database 855 or retrieved for use with the test sample. For example, the model is a trained cancer classifier for determining cancer predictions for a test sample using feature vectors derived from highly informative fragments. The training and use of cancer classifiers will be discussed further in conjunction with Section IV, Cancer Classifiers for Determining Cancer. The analysis system 800 may train one or more models 850 and store various trained parameters in the parameter database 865. The analysis system 800 stores the models 850 along with their functions in the model database 855.

[0159] During inference, the score engine 860 returns an output using one or more models 850. The score engine 860 accesses the models 850 in the model database 855, along with the trained parameters from the parameter database 865. According to each model, the score engine receives the appropriate input for the model and calculates the output based on the received input, parameters, and the function of each model that correlates the input and output. In some use cases, the score engine 860 further calculates metrics that correlate with the reliability of the calculated output from the model. In other use cases, the score engine 860 calculates other intermediate values ​​for use with the model.

[0160] II. Sample Sequencing and Processing II.A. Generation of Methylation State Vectors for DNA Fragments Figure 2A is an exemplary flowchart illustrating a process 200 for sequencing a cfDNA fragment to obtain a methylation state vector, according to one or more embodiments. To analyze DNA methylation, the analysis system first obtains a sample from an individual containing multiple cfDNA molecules 210. In additional embodiments, process 200 can be applied to sequence other types of DNA molecules. Process 200 is one embodiment of the sample sequencing 120 in Figure 1.

[0161] From the sample, the analytical system can isolate each cfDNA molecule.210 The cfDNA molecules can be processed to convert unmethylated cytosine to uracil.220 In one embodiment, the method uses bisulfite treatment of DNA to convert unmethylated cytosine to uracil without converting methylated cytosine. For example, commercially available kits such as EZ DNA Methylation®-Gold, EZ DNA Methylation®-Direct, or an EZ DNA Methylation®-Lightning kit (available from Zymo Research Corp (Irvine, CA)) are used for bisulfite conversion. In another embodiment, the conversion from unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for the conversion of unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0162] A sequencing library can be prepared from the converted cfDNA molecule.230 During library preparation, a unique molecular identifier (UMI) may be added to a nucleic acid molecule (e.g., a DNA molecule) via adapter ligation. The UMI may be a short nucleic acid sequence (e.g., 4-10 base pairs) added to the end of a DNA fragment (e.g., a DNA molecule fragmented by physical shear, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. The UMI may be a degenerate base pair that functions as a unique tag that can be used to identify sequence reads originating from a particular DNA fragment. During PCR amplification after adapter ligation, the UMI may be replicated along with the bound DNA fragment. This may provide a method for identifying sequence reads originating from the same original fragment in downstream analysis.

[0163] Optionally, a sequencing library may be enriched with cfDNA molecules or genomic regions that are highly informative about cancer status using multiple hybridization probes.235 Hybridization probes are short oligonucleotides that hybridize to specific cfDNA molecules or target regions, and can enrich those fragments or regions for subsequent sequencing and analysis. Using hybridization probes, researchers can perform targeted, high-depth analysis of a specific set of CpG sites of interest. Hybridization probes can be arranged across one or more target sequences with coverage of 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x or more. For example, hybridization probes arranged with 2x coverage include duplicate probes so that each portion of the target sequence hybridizes to two independent probes. Hybridization probes can be arranged across one or more target sequences with coverage of less than 1x.

[0164] In one embodiment, a hybridization probe is designed to enrich DNA molecules that have been treated (e.g., using bisulfite) for conversion from unmethylated cytosine to uracil. During enrichment, the hybridization probe (also referred to herein as the “probe”) can be used to target and pull down highly informative nucleic acid fragments about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer classification or tissue of origin). The probe may be designed to anneal to (or hybridize to) a target (complementary) strand of DNA. The target strand may be a “plus” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “minus” strand. The probe may range in length from 10, 100, or 1000 base pairs. The probe may be designed based on a methylation site panel. The probe may be designed based on a panel of target genes to analyze specific mutations or target regions of a genome (e.g., the human or another organism’s genome) suspected to correspond to a particular cancer or other type of disease. In addition, the probe may cover the overlapping portion of the target region.

[0165] Once prepared, a sequencing library or a portion thereof can be sequenced to obtain multiple sequence reads.240 Sequence reads may be in a computer-readable digital format for processing and interpretation by computer software. Sequence reads may be aligned to a reference genome to determine alignment position information. Alignment position information may indicate the start and end positions of a region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. Alignment position information may also include the sequence read length, which can be determined from the start and end positions. The region in the reference genome may be associated with a gene or a segment of a gene. Sequence reads may consist of read pairs, denoted as R1 and R2. For example, the first read R1 may be sequenced from the first end of a nucleic acid fragment, while the second read R2 may be sequenced from the second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 may be aligned in accordance with the nucleotide bases of the reference genome (e.g., in opposite directions). The alignment position information derived from the read pair R1 and R2 may include the start position in the reference genome corresponding to the end of the first read (e.g., R1) and the end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome can represent possible locations of nucleic acid fragments within the corresponding reference genome. Output files in SAM (Sequence Alignment Map) format or BAM (Binary) format are generated and can be output for further analysis, such as methylation status determination.

[0166] From the sequence reads, the analysis system determines the location and methylation status of each CpG site based on alignment with the reference genome.250 The analysis system generates a methylation status vector for each fragment, specifying the location of the reference genome fragment (e.g., identified by the location of the first CpG site in each fragment or another similar metric), the number of CpG sites in the fragment, and the methylation status of each CpG site in the fragment, whether methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or undeterminable (e.g., denoted as I).260 Observed states can be methylated and unmethylated, while unobserved states are undeterminable. Undeterminable methylation states may result from sequencing errors and / or mismatches between the methylation states of the complementary strands of the DNA fragments. The methylation status vector may be stored in temporary or persistent computer memory for later use and processing. Furthermore, the analysis system may remove replicated reads or replicated methylation status vectors from a single sample. The analysis system can determine that certain fragments having one or more CpG sites have an undeterminable methylation state exceeding a threshold number or percentage. Such fragments may be excluded or selectively included, and a model can be constructed to explain these undeterminable methylation states.

[0167] Figure 2B is an illustrative diagram of process 200 of Figure 2A for sequencing a fragment of a cfDNA molecule to obtain a methylation state vector, according to one or more embodiments. As an example, the analysis system receives a cfDNA molecule 212 containing three CpG sites in this example. As shown, the first and third CpG sites of cfDNA molecule 212 are methylated 214. During processing step 220, cfDNA molecule 212 is transformed to produce a transformed cfDNA molecule 222. During processing 220, the second CpG site, which is in an unmethylated state, has its cytosine converted to uracil. However, the first and third CpG sites were not transformed.

[0168] After conversion, a sequencing library 230 is prepared and sequenced 240, generating sequence reads 242. The analysis system aligns sequence reads 242 to a reference genome 244 250. The reference genome 244 provides context regarding the location of the fragment cfDNA originating in the human genome. In this simplified example, the analysis system aligns sequence reads 242 so that three CpG sites correlate with CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of explanation) 250. Thus, the analysis system can generate information on both the methylation status of all CpG sites on the cfDNA molecule 212 and the location in the human genome to which the CpG sites map. As shown, the methylated CpG sites on sequence read 242 are read as cytosine. In this example, cytosine appears in sequence read 242 only at the first and third CpG sites, which allows us to infer that the first and third CpG sites in the original cfDNA molecule are methylated. On the other hand, the second CpG site can be read as thymine (U is converted to T during the sequencing process), and therefore it can be inferred that the second CpG site is unmethylated in the original cfDNA molecule. Using these two pieces of information, methylation state and location, the analysis system generates a methylation state vector 252 of the fragment cfDNA 212 260. In this example, the resulting methylation state vector 252 is, <M 23 , U 24 M 25 > represents a methylated CpG site, U represents a non-methylated CpG site, and the subscript number corresponds to the location of each CpG site in the reference genome.

[0169] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in biological samples. These sequencing methods include, but are not limited to, high-throughput sequencing systems such as the Roche 454 platform, Applied Biosystems SOLID platform, and Helicos True Single Molecule DNA sequencing technology; hybridization sequencing platforms from Affymetrix Inc.; single-molecule, real-time (SMRT) technology from Pacific Biosciences; synthetic sequencing platforms from 454 Life Sciences, Illumina / Solexa, and Helicos Biosciences; and ligation sequencing platforms from Applied Biosystems; and any other form of sequencing that can be used to obtain a number of sequence reads measured from nucleic acids (e.g., cell-free nucleic acids). ION TORRENT technology and nanopore sequencing from Life Technologies can also be used to obtain sequence reads from biological samples (e.g., cell-free nucleic acids). Synthetic sequencing and reversible terminator-based sequencing (e.g., Illumina Genome Analyzer, Genome Analyzer II, HISEQ 2000, HISEQ 4500 (Illumina, San Diego, Calif.)) can be used to obtain sequence reads from cell-free nucleic acids obtained from training biological samples to form genotype datasets. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one example of this type of sequencing technique, a flow cell containing an optically transparent slide with eight individual lanes on a surface to which oligonucleotide anchors (e.g., adapter primers) are bound is used. Cell-free nucleic acid samples may contain signals or tags to facilitate detection.The acquisition of sequence reads derived from cell-free nucleic acids obtained from biological samples may include obtaining quantitative information on signals or tags through various techniques such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarrays, mass spectrometry, cell fluorescence analysis, fluorescence microscopy, confocal laser microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and combinations thereof.

[0170] One or more sequencing methods may include whole-genome sequencing assays. Whole-genome sequencing assays may include physical assays that generate sequence reads of the whole genome or a substantial portion of the whole genome, which can be used to determine large variations such as copy number variation or copy number anomalies. Such physical assays may use whole-genome sequencing technology or whole-exome sequencing technology. Whole-genome sequencing assays may have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the entire genome under test. In some embodiments, the sequencing depth is approximately 30,000x. One or more sequencing methods may include targeted panel sequencing assays. Targeted panel sequencing assays may have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x for a targeted gene panel. The target gene panel may include 450 to 500 genes. The target gene panel may include genes in the range of 500 ± 5, 500 ± 10, or 500 ± 25.

[0171] One or more sequencing methods may include paired-end sequencing. One or more sequencing methods may generate multiple sequence reads. Multiple sequence reads may have characteristic lengths in the range of 10–700, 50–400, or 100–300. One or more sequencing methods may include methylation sequencing assays. Methylation sequencing may be i) whole-genome methylation sequencing, or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, methylation sequencing is whole-genome bisulfite sequencing (e.g., WGBS). Methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes targeting the most informative regions of methylome, a proprietary methylation database, and previous prototype whole-genome and targeted sequencing assays.

[0172] Methylation sequencing can detect one or more 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) in each nucleic acid methylation fragment. Methylation sequencing may involve the conversion of one or more unmethylated cytosines or one or more methylated cytosines in each nucleic acid methylation fragment to one or more corresponding uracils. One or more uracils may be detected as one or more corresponding thymines during methylation sequencing. The conversion of one or more unmethylated cytosines or one or more methylated cytosines may include chemical conversion, enzymatic conversion, or a combination thereof.

[0173] For example, bisulfite conversion involves converting cytosine to uracil while keeping methylated cytosine (e.g., 5-methylcytosine or 5-mC) intact. In some DNAs, approximately 95% of cytosine may not be methylated in the DNA, and the resulting DNA fragment may contain a large amount of uracil represented as thymine. Nucleic acids can be treated using enzymatic conversion processes before sequencing, and this can be carried out in various ways. An example of bisulfite-free conversion is bisulfite-free and base-resolution sequencing, TET-assisted pyridine borane sequencing (TAPS), for the non-destructive and direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosine. The methylation status of the CpG sites in the corresponding multiple CpG sites within each nucleic acid methylation fragment is such that methylation may occur when methylation sequencing determines that the CpG site is methylated, and may not occur when methylation sequencing determines that the CpG site is not methylated.

[0174] Methylation sequencing assays (e.g., WGBS and / or targeted methylation sequencing) may have average sequencing depths of approximately 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. Methylation sequencing may have sequencing depths greater than 30,000x, for example, at least 40,000x or 50,000x. Whole-genome bisulfite sequencing methods may have average sequencing depths of 20x to 50x, and targeted methylation sequencing methods may have average effective depths of 100x to 1,000x, although the effective depth may be equivalent whole-genome bisulfite sequencing coverage to obtain the same number of sequence reads obtained by targeted methylation sequencing.

[0175] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, for example, U.S. Patent Application No. 16 / 352,602, filed March 13, 2019, entitled "Methylation Fragment Anomaly Detection," and U.S. Patent Application No. 16 / 719,902, filed December 18, 2019, entitled "Systems and Methods for Estimating Cell Source Fractions Using Methylation Information," respectively, which are incorporated herein by reference. Fragment methylation patterns can be obtained using the methods disclosed herein and / or other methods for methylation sequencing, including any modifications, substitutions, or combinations thereof. For example, one or more methylation state vectors can be identified using methylation sequencing in accordance with any of the techniques described in U.S. Patent Application No. 16 / 352,602, filed March 13, 2019, entitled "Anomalous Fragment Detection and Classification," each of which is incorporated herein by reference, or disclosed in U.S. Patent Application No. 15 / 931,022, filed May 13, 2020, entitled "Model-Based Featuration and Classification."

[0176] Multiple nucleic acid methylation fragments can be obtained using nucleic acid methylation sequencing and one or more methylation state vectors obtained. Each corresponding set of multiple nucleic acid methylation fragments (e.g., for each genotype dataset) may contain more than 100 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding set of multiple nucleic acid methylation fragments may include more than 1,000, more than 5,000, more than 10,000, more than 20,000, or more than 30,000 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding set of multiple nucleic acid methylation fragments may be between 10,000 and 50,000 nucleic acid methylation fragments. The corresponding set of nucleic acid methylation fragments may include more than 1,000, 10,000, 100,000, 1,000,000, 10,000,000, 100,000, 500,000, 1,0

[0177] Further details regarding methods for sequencing nucleic acids and methylation sequencing data are disclosed in U.S. Patent Application No. 17 / 191,914, filed on March 4, 2021, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," which is incorporated herein by reference.

[0178] III. Detection of WBC contamination The analytical system implements WBC contamination detection to detect WBC contamination in samples used in cancer classification workflows. Specifically, the analytical system aims to prevent samples contaminated with WBC efflux DNA from being obtained for downstream analysis, as WBC efflux DNA can interfere with cancer detection during analysis in plasma combined with cfDNA, potentially leading to false positive calls for cancer. WBC contamination detection allows the analytical system to determine whether a sample is contaminated with WBC efflux DNA and may further determine the level of contamination. Such screening ensures that training data used to train downstream models is not distorted by WBC contamination. Screening can also reduce false positive predictions by refraining from using contaminated samples for further analysis. Upon detection of a contaminated sample, further corrective actions can be taken to address the contamination. Removal of the contamination source improves the physical assay process and sample handling.

[0179] This specification describes three different methods for detecting WBC contamination. The first methodology, called "coverage-based WBC contamination detection," aims to assess whether a sample has abnormal coverage in a set of genomic loci, indicating WBC contamination. The second methodology, called "methylation-based WBC contamination detection," aims to deconvolve the proportion of tissue types based on methylation features derived from sequence reads in the sample. Based on the deconvolved proportion of tissue types, the analytical system may determine the distribution of tissue type proportions from a set of purified cfDNA samples to assess the likelihood that the test sample is contaminated. The purified cfDNA samples are purified to remove any non-cfDNA fragments, such as WBC-derived fragments. The third methodology, called "quantitative coverage-based WBC contamination detection," aims to assess the amount of WBC-extracted DNA present in the sample based on maximizing the likelihood of observing coverage across a set of genomic locus features. The principles explained throughout these methodologies can be combined in various ways to form hybrid methodologies for evaluating WBC contamination.

[0180] Each sample includes sequence reads derived from sequencing of DNA fragments contained in each physically collected biological sample. Sequence reads generally represent, for example, the amino acid sequence of the DNA fragment in the biological sample. Sequence reads may include methylation information, as illustrated in Figures 2A and 2B. In some embodiments, sequence reads may be derived from whole-genome bisulfite sequencing (WGBS) and / or targeted sequencing. In WGBS, there are sequence reads for each DNA fragment (or DNA molecule) contained in the biological sample. In WGBS, sequence reads are mapped to targeted genomic regions (or genomic loci) pulled down from a targeted probe. Otherwise, sequence reads may include other types of sequencing information, such as single nucleotide polymorphisms, insertions or deletions, or single tandem repeats.

[0181] III.A. Coverage-based detection of WBC contamination Figures 3A–3C include flowcharts illustrating a first methodology for coverage-based WBC contamination detection. Specifically, Figure 3A shows a flowchart illustrating feature identification for coverage-based WBC contamination detection in one or more embodiments. Figure 3B shows a flowchart illustrating training a contamination model for predicting coverage-based WBC contamination across a set of genomic locus features in one or more embodiments. Figure 3C shows a flowchart illustrating WBC contamination prediction using the contamination model in one or more embodiments. The analysis system is described as performing each of the processes described in the flowcharts. Nevertheless, in other embodiments, various computing systems may perform each process.

[0182] Figure 3A shows a flowchart illustrating feature identification 300 for coverage-based WBC contamination detection according to one or more embodiments.

[0183] The analysis system obtains a first set of cfDNA samples from a first cohort of subjects and a second set of WBC samples from the same first cohort of subjects.305 For example, each subject in the first cohort provides at least one cfDNA sample and at least one WBC sample. The cfDNA sample contains sequence reads relating to a cfDNA fragment. The WBC sample contains sequence reads relating to a WBC efflux DNA fragment.

[0184] The analysis system determines, for each sample, the coverage of sequence reads overlapping with each genomic locus in an initial set of genomic loci.310 The initial set of genomic loci may include up to all genomic regions, e.g., all targeted genomic regions in a targeted sequencing embodiment or all genomic regions in a WGBS embodiment. Coverage can be counted by identifying unique sequence reads that partially, substantially, or completely overlap with the genomic regions. In some embodiments, the genomic region includes at least one CpG site. In other embodiments, the genomic region extends to a set of CpG sites. In yet another embodiment, the genomic region is at least 50 bp, 100 bp, 150 bp, 200 bp, or 250 bp in length. From the initial set of genomic regions, the analysis system can prune genomic regions where cfDNA coverage across the sample is variable. The analysis system can evaluate the statistical variance of coverage across the cfDNA sample to determine whether a particular genomic region has excessive variability beyond, for example, statistical variance or distribution spread. The analysis system can further evaluate whether the coverage follows a normal distribution, for example, whether pruning or removal of genomic regions that have a non-normal distribution is appropriate.

[0185] As an example, for the cfDNA sample of subject A, the analysis system determines the coverage of genomic locus 1 based on sequence reads in the cfDNA sample, the coverage of genomic locus 2 based on sequence reads in the cfDNA sample, etc., through the remaining genomic loci in the initial set. Similarly, for the WBC sample of subject A, the analysis system determines the coverage of genomic locus 1 based on sequence reads in the WBC sample, the coverage of genomic locus 2 based on sequence reads in the WBC sample, etc., through the remaining genomic loci in the initial set.

[0186] The analysis system can further normalize the coverage of each sample. For example, the analysis system determines the sequencing depth based on the sequence reads in the sample. The sequencing depth allows the analysis system to normalize each coverage by dividing by the sequencing depth. Normalizing the coverage ensures that differential sequencing depths between samples do not distort the analysis.

[0187] The analysis system determines the coverage ratio of the WBC sample to the cfDNA sample for each genomic locus.315 The analysis system can calculate the coverage ratio for each genomic locus between the WBC sample and the cfDNA sample for each subject. Following the above example, for subject A, the analysis system can calculate the coverage ratio for genomic locus 1 using the remaining genomic loci in the initial set by dividing the coverage of genomic locus 1 from the WBC sample by the coverage of genomic locus 1 from the cfDNA sample. If each subject has its own coverage ratio across genomic loci, the analysis system can determine the total coverage ratio of the WBC sample to the cfDNA sample by averaging the coverage ratios.

[0188] The analysis system determines a set of features for genomic loci whose coverage ratio exceeds a threshold, for use in evaluating WBC contamination.320 Genomic loci in the feature set may also be called "WBC contamination markers" (or more generally, "contamination markers"). The analysis system determines the feature set using the threshold ratio. The analysis system can further rank the genomic loci based on the coverage ratio and select the top number from the ranking.

[0189] Figure 3B shows a flowchart illustrating the training of a contamination model for predicting coverage-based WBC contamination across feature sets of genomic loci in one or more embodiments.

[0190] The analysis system obtains a set of cfDNA samples from non-cancer subjects, comprising a first subset and a second subset.335 The cfDNA samples may be purified to ensure the absence of confusing WBC efflux DNA signals and may be further purified to ensure the absence of blood disorder signals. The set of cfDNA samples from non-cancer subjects may be divided into a first subset and a second subset. The first subset can be used to train a contamination model and predict contamination metrics for the samples, and the second subset is used to determine a contamination threshold for calling WBC contamination from the contamination metrics. In some embodiments, the subsets may be the same, even if they overlap.

[0191] The analysis system determines the coverage of sequence reads that overlap with each genomic locus in the feature set for each sample in the first and second subsets.340 Coverage can be counted by identifying unique sequence reads that partially, substantially, or completely overlap with genomic regions. Coverage may be further normalized, for example, according to the sequencing depth of the sample.

[0192] The analysis system generates a coverage-based distribution for each genomic locus in the feature set, based on the coverage of a first subset of samples.345 The distribution for each genomic locus can be explained by its mean and standard deviation. This distribution is useful for determining the likelihood of observing coverage at a genomic locus in the test sample. At the time of generation, there is a coverage distribution for each genomic locus based on samples from non-cancer subjects.

[0193] The analysis system determines the statistical likelihood of observing coverage at each genomic locus for each sample in the second subset, based on the coverage distribution for that locus.350 For example, with respect to sample B in the second subset, the analysis system determines the statistical likelihood of observing coverage at genomic locus 1, based on the coverage distribution for that locus. In some embodiments, the statistical likelihood of coverage is a z-score based on the coverage distribution. The z-score represents the number of standard deviations away from the mean of the distribution. In other embodiments, the statistical likelihood of coverage is a p-value based on the coverage distribution, representing the likelihood of observing the sample's coverage or more at that genomic locus. The p-value is determined by the proportion of the distribution of coverage greater than the sample's coverage.

[0194] The analysis system generates contamination metrics for each sample in the second subset by combining the statistical likelihoods between genomic loci in the feature set.355 In embodiments with z-scores, the analysis system can generate contamination metrics for a sample by summing the z-scores across genomic loci in the feature set. In some embodiments, the analysis system can sum the absolute values ​​of the z-scores. In some embodiments, the analysis system can generate contamination metrics by summing the absolute values ​​of z-scores that exceed a threshold. In embodiments with p-value scores, the analysis system can generate contamination metrics as a truncated p-value product of p-values ​​across genomic loci. The truncated p-value product can be calculated by combining p-values ​​below a certain specific cutoff value.

[0195] The analysis system determines the distribution of contamination metrics for a second subset of the sample.360 The analysis system can generate a distribution using the contamination metrics calculated for the second subset of the sample. The distribution can be expressed by its mean and standard deviation.

[0196] The analysis system determines the contamination threshold based on the distribution of contamination metrics.365 As mentioned above, a second subset of the sample comes from non-cancer subjects (e.g., otherwise healthy subjects). Therefore, the distribution represents a normal distribution of non-contaminated samples from non-cancer subjects. Using this distribution, the analysis system can set specificity thresholds such as 95%, 96%, 97%, 98%, 99%, and 99.5%. Based on these specificity thresholds, the analysis system can determine the contamination threshold to use when calling whether WBC contamination is present. The contamination threshold divides the distribution such that having contamination metrics above the contamination threshold indicates that the sample is likely to be contaminated with WBCs, as the probability of naturally observing such contamination metrics is very low. In other words, the proportion of the distribution below the contamination threshold is equal to the specificity threshold. A contamination model can be generated and input the coverage across the feature set of the sample's genomic loci to determine whether the sample is contaminated with WBCs.

[0197] The analysis system can further evaluate the predictability of the feature set of the contamination model against the null distribution of other candidate sets. The analysis system can identify genomic loci from the initial set with equal coverage between cfDNA samples and WBC samples. From the set of genomic loci with equal coverage, the analysis system can randomly select candidate sets to compare predictability. Each candidate set of genomic loci may be the same size as the feature set. For example, if the feature set contains 50 genomic loci (or contamination markers), the candidate set may also contain 50 genomic loci (or contamination markers). The analysis system can similarly train alternative contamination models according to model training 330. Using the contamination models trained with the feature sets and alternative contamination models, the analysis system can use a validation set of samples containing samples with known WBC contamination (e.g., in silico or titrated from WBC samples). The analysis system can evaluate how many contaminated samples were correctly identified using a contamination model trained with a feature set, as well as alternative contamination models.

[0198] Figure 3C shows a flowchart illustrating WBC contamination prediction 370 using a contamination model in one or more embodiments. In one or more embodiments, the contamination model includes various distributions generated in the model training 330 in Figure 3B. For example, the contamination model includes a coverage distribution and a contamination metric distribution. The contamination model may further include a contamination threshold for calling WBC contamination in a sample.

[0199] The analysis system obtains a test sample from the subject of study.375 The test sample may contain sequence reads sequenced from a biological sample taken from the subject of study. The test sample may be used to validate a contamination model or may be derived from a subject undergoing cancer classification. The biological sample may be a blood sample processed to isolate solid clumps containing plasma, as well as red blood cells and other microparticles in the blood, from the WBC buffy coat. The sequence reads may correspond to cfDNA fragments in the plasma. The analysis system applies a contamination model (e.g., one trained by model training in Figure 3B) to the test sample to evaluate whether the sample contains WBC contamination.

[0200] The analysis system determines the coverage of sequence reads that overlap with each genomic locus in the feature set.380 Coverage can be counted by identifying unique sequence reads that partially, substantially, or completely overlap with genomic regions. Coverage may be further normalized, for example, according to the sequencing depth of the sample.

[0201] The analysis system determines the statistical likelihood of observing coverage at each genomic locus based on the coverage distribution for that locus.385 In some embodiments, the statistical likelihood of coverage is a z-score based on the coverage distribution. In other embodiments, the statistical likelihood of coverage is a p-value based on the coverage distribution, indicating the likelihood of observing sample coverage or exceeding coverage at that genomic locus.

[0202] The analysis system generates contamination metrics for test samples by combining statistical likelihoods across genomic loci.390 In embodiments with z-scores, the analysis system can generate contamination metrics for a sample by summing the z-scores across genomic loci in the feature set. In some embodiments, the analysis system can sum the absolute values ​​of the z-scores. In some embodiments, the analysis system can generate contamination metrics by summing the absolute values ​​of z-scores that exceed a threshold. In embodiments with p-value scores, the analysis system can generate contamination metrics as a truncated p-value product of p-values ​​across genomic loci. The truncated p-value product can be calculated by combining p-values ​​below a certain cutoff value.

[0203] The analysis system determines whether a test sample has WBC contamination if the contamination metric for the test sample is above the contamination threshold.395 The contamination threshold is a value identified based on the distribution of contamination metrics to ensure a certain degree of specificity to avoid false positive calls for WBC contamination.

[0204] If it is determined that a test sample is contaminated with WBCs, the analytical system may implement corrective actions. Firstly, the analytical system may provide the healthcare provider with a notification indicating the detection of WBC contamination in the test sample. The notification may further encourage the healthcare provider to collect subsequent biological samples for sequencing and analysis in the cancer classification workflow. The notification may further inform whether certain clinical supplies, specific clinicians, or sample handling techniques are more likely to result in WBC contamination. Secondly, the analytical system may discourage the use of the test sample for further analysis in the cancer classification workflow to avoid false-positive calls of cancer due to confusion of WBC-shedding DNA. Discouraging the use of samples from downstream analyses may involve filtering of training samples, improving the training of cancer classification models.

[0205] III.B. Methylation-based detection of WBC contamination Figures 4A to 4C include flowcharts illustrating a second methodology for methylation-based WBC contamination detection. In particular, Figure 4A shows a flowchart illustrating the training of a deconvolution model for methylation-based WBC contamination detection 400 by one or more embodiments. Figure 4B shows a flowchart illustrating the training of a contamination model for predicting WBC contamination based on the proportion of predicted tissue types 440 by one or more embodiments. Figure 4C shows a flowchart illustrating WBC contamination prediction 470 using a contamination model by one or more embodiments. The analysis system is described as performing each of the processes described in the flowcharts. Nevertheless, in other embodiments, various computing systems may perform each process.

[0206] Figure 4A shows a flowchart illustrating the training of a deconvolution model for methylation-based WBC contamination detection according to one or more embodiments. The deconvolution model can be trained to take methylation information as input and output the percentage of tissue types.

[0207] The analysis system obtains a set of samples from different tissue types.405 Each set of samples may be labeled by a tissue type from a different tissue type. In some embodiments, the set of samples may be further separated into subsets based on tissue subtypes. A sample from one tissue type generally comes from a biological sample containing DNA fragments derived from cells of that tissue type. For example, the biological sample may be a biopsy sample taken from a particular tissue. The tissue type may also include a cell type.

[0208] In some embodiments, the set of tissue types is selected as a combination of the following tissue types: erythrocyte progenitor type, megakaryocyte type, monocyte type, granulocyte type, lymphocyte type, epithelial cell type, vascular endothelial cell type, hepatocyte type, adipocyte type, endocrine cell type, and myocyte type. In other embodiments, the set of tissue types is selected as a combination of the following tissue types: circulating bone marrow tissue types (e.g., including erythrocyte progenitor type and megakaryocyte type), non-circulating bone marrow tissue types (e.g., including monocyte type and granulocyte type), lymphocyte type, epithelial cell type, vascular endothelial cell type, and hepatocyte type.

[0209] The analysis system determines methylation features at each genomic locus for each sample, based on the sequence reads.410 A genomic locus may contain one or more methylation sites, e.g., CpG sites. A genomic locus may be an initial set of genomic loci, e.g., covered by the sequencing process, and may be up to all genomic loci. Sequence reads may further contain methylation information, as determined, for example, by Figures 2A and 2B. The methylation features characterize the methylation state of one or more sequence reads that overlap with a genomic locus. In some embodiments, the analysis system determines multiple methylation features at one or more genomic loci. In one example, the methylation features may be the methylation density at a genomic locus. In another example, the methylation features may be the count (or percentage) of sequence reads that overlap with a genomic locus that is either highly methylated or highly unmethylated. In a third embodiment, the methylation features may be the count (or percentage) of sequence reads that have a specific methylation variant at a genomic locus.

[0210] The analysis system generates a feature vector for each sample based on methylation features across genomic loci 415. The feature vector represents the methylation features as determined in step 410. For example, the feature vector may enumerate all methylation features as elements. In other examples, the analysis system may combine various methylation features to reduce the dimensionality of the feature vector.

[0211] The analysis system trains a deconvolution model to predict the proportion of tissue types based on feature vectors from a sample.420 The deconvolution model can be a machine learning model. Generally, the deconvolution model is trained to predict the tissue type of a sample by taking feature vectors from the sample as input. Based on the accuracy of the prediction, the analysis system can adjust the weights or other parameters in the deconvolution model to improve the accuracy of the prediction. In some embodiments, the analysis system may train the deconvolution model in multiple folds. In each fold, the analysis system adjusts the weights and / or parameters in the deconvolution model using a subset of training samples. The trained deconvolution model can take feature vectors representing methylation information of sequence reads as input and output predictions of the proportion of tissue types. The proportion of each tissue type is the percentage of the tissue type's contribution to the DNA fragments (or sequence reads) present in the sample. For example, for a set of six tissue types, the deconvolution model is configured to output six tissue type proportions and one tissue type proportion for each tissue type. The sum of the tissue type proportions may correspond to 100% or 1. The analysis system can further improve and optimize the deconvolution model.

[0212] The analysis system can identify feature sets of genomic loci based on the information gain obtained from the deconvolution model.425 The analysis system can evaluate the information gain by evaluating the entropy of each feature (or methylation feature) in the feature vector. For example, mutual information can provide a measure of the interdependence between two conditions of interest (e.g., tissue types) sampled simultaneously. The mutual information score may represent the discriminative power of a methylation feature to distinguish between two tissue types. Information gain can be evaluated during the training of the deconvolution model, for example, after each training epoch. The analysis system can rank features based on their information gain. The analysis system can identify feature sets as features with an information gain above an information gain threshold. In other embodiments, the analysis system can identify feature sets as the top number of features ranked based on information gain. The analysis system can select feature sets based on other selection criteria. For example, the analysis system can identify features that are sufficiently spread across the entire genome. As another example, the analysis system can assess whether the features are stable, for example, whether they are normally distributed across the entire population. A set of features may represent one or more methylation features at one or more genomic loci.

[0213] The analysis system can retrain the deconvolution model using the updated feature vector based on the identified feature set.430 The analysis system can update the feature vector, for example, to retain features in the feature set. For example, feature selection may yield the set of genomic loci that are most discriminative when predicting tissue type, and the analysis system can update the sample's feature vector to retain only the methylation features for that feature set of genomic loci. The analysis system can then retrain the deconvolution model using the updated feature vector. The analysis system can further use a holdout set of samples to evaluate the predictive accuracy of the deconvolution model.

[0214] Figure 4B shows a flowchart illustrating the training of a contamination model 440 for predicting WBC contamination based on the proportion of predicted tissue types, according to one or more embodiments.

[0215] The analytical system obtains a set of cfDNA samples from non-cancer subjects.445 Each cfDNA sample can be purified to ensure that the DNA fragments and corresponding sequence reads are not contaminated with WBC-extracted DNA.

[0216] The analysis system determines methylation features at each genomic locus in the feature set based on sequence reads for each sample.450 The analysis system can determine multiple features for each genomic locus. The feature set for a genomic locus can be determined, for example, by evaluating the information gain, as explained in Figure 4A.

[0217] The analysis system generates a feature vector for each sample based on methylation features across a set of feature characteristics of genomic loci.455 The feature vector may include methylation features across a set of feature characteristics of genomic loci, or may include several combination representations of methylation features.

[0218] The analysis system applies a deconvolution model to the feature vectors of each sample to predict the proportion of tissue types.460 The deconvolution model can be trained according to the training described in Figure 4A.400 The proportion of each tissue type is the predicted contribution from the DNA fragments (or their sequence reads) in the sample. The proportion of tissue types can be represented by a vector.

[0219] The analysis system generates a distribution of tissue type proportions in a sample.465 The distribution is a multivariate distribution in a multivariate vector space. The number of variables corresponds to the number of predicted tissue types. The distribution is formed by the predicted proportions of tissue types in the sample. The distribution can be represented by the mean proportion of tissue types as a representative value of the proportions of tissue types across the sample, and further by the principal axes of the vector space. This distribution represents the null distribution of cfDNA samples from non-cancer populations. The contamination model implements the distribution and calculates a p-value representing the likelihood of observing the predicted proportions of tissue types in the test sample in the null distribution. Except for very low likelihoods, it is considered to be within the predicted range for cfDNA samples without WBC contamination. Very low likelihoods are considered outliers of the distribution and indicate WBC contamination.

[0220] Figure 4C shows a flowchart illustrating WBC contamination prediction 470 using a contamination model in one or more embodiments.

[0221] The analysis system obtains a test sample from the subject of test.475 The test sample may be used to validate a contamination model or may be derived from a subject undergoing cancer classification. The test sample contains sequence reads relating to DNA fragments in a biological sample taken from the subject of test. In one or more embodiments, the test sample undergoes cancer classification to determine whether the cfDNA fragment in the biological sample indicates the presence or possibility of cancer. The analysis system applies a contamination model to predict whether the test sample is contaminated with WBCs and to prevent confusion between cancer classification and WBC exudate DNA.

[0222] The analysis system determines methylation features at each genomic locus in a feature set based on sequence reads for each sample.480 Methylation features may include, for example, methylation density, the count or percentage of highly methylated sequence reads, the count or percentage of highly unmethylated sequence reads, or the count or percentage of sequence reads with a specific methylation variant.

[0223] The analysis system generates a feature vector for the test sample based on methylation features.485 The feature vector may enumerate all methylation features, or it may be a combination of one or more methylation features.

[0224] The analysis system applies a deconvolution model to feature vectors to predict the proportion of tissue types in the test sample.490 The proportion of tissue types represents the contribution rate of multiple tissue types derived from DNA fragments (or corresponding sequence reads) in the biological sample.

[0225] The analysis system generates sample contamination metrics based on the distance of the test sample to the distribution of tissue type proportions.495 In one or more embodiments, the distance may be calculated based on the mean tissue type proportions and one or more principal components that describe the distribution of tissue type proportions. In one or more embodiments, the distance is the Mahalanobis distance:

number

number

[0226] The analysis system determines whether a sample has WBC contamination based on the distribution of contamination metrics.497 In one or more embodiments, the contamination metric is a p-value indicating the likelihood of observing a predicted proportion of a particular tissue type in a cfDNA sample from a non-cancer population. In such embodiments, the p-value is compared to a p-value threshold. If the p-value is less than the p-value threshold, the analysis system calls the test sample contaminated with WBCs. If the p-value is greater than the p-value threshold, the analysis system calls the test sample not contaminated with WBCs.

[0227] If it is determined that a test sample is contaminated with WBCs, the analytical system may take corrective actions. Firstly, the analytical system may provide the healthcare provider with a notification indicating the detection of WBC contamination in the test sample. The notification may further recommend that the healthcare provider collect subsequent biological samples for sequencing and analysis in the cancer classification workflow. The notification may further inform whether certain clinical supplies, specific clinicians, or sample handling techniques are more likely to result in WBC contamination. Secondly, the analytical system may discourage the use of the test sample for further analysis in the cancer classification workflow to avoid false-positive calls of cancer due to confusion of WBC-shedding DNA. Discouraging the use of samples from downstream analyses may involve filtering of training samples, improving the training of cancer classification models.

[0228] III.C. Quantitative coverage-based detection of WBC contamination Figures 5A and 5B include flowcharts illustrating a third methodology for quantitative coverage-based WBC contamination detection. In particular, Figure 5A shows a flowchart illustrating feature identification and training of a contamination model for quantitative coverage-based WBC contamination detection in one or more embodiments. Figure 5B shows a flowchart illustrating WBC contamination prediction using the contamination model in one or more embodiments. The analysis system is described as performing each of the processes described in the flowcharts. Nevertheless, in other embodiments, various computing systems may perform each process.

[0229] Figure 5A shows a flowchart illustrating feature recognition and training of a contamination model for quantitative coverage-based WBC contamination detection according to one or more embodiments.

[0230] The analysis system obtains a first set of cfDNA samples from the target cohort and a second set of WBC gDNA samples from the target cohort.505 The cfDNA samples may be purified to ensure that only cfDNA fragments constitute a biological sample from which the sequence reads originate. The WBC gDNA samples may be purified to ensure that only WBC genomic DNA fragments constitute a biological sample from which the sequence reads originate. Each subject in the cohort may provide at least one cfDNA sample and at least one WBC gDNA sample.

[0231] The analysis system determines the average coverage for each sample.510 In some embodiments, the analysis system determines the average coverage across all genomic loci or sequenced markers. Genomic loci may include binary markers and semi-binary markers. Binary markers evaluate whether a sequence read is one of two states at a genomic locus. Semi-binary markers evaluate whether a sequence read exists that covers one of several states at a genomic locus. In some embodiments, the analysis system determines the average coverage as a representative value across binary markers.

[0232] The analysis system determines the normalized coverage of each locus (or marker) in the initial set of genomic loci for each sample.515 Each sample is normalized separately by its average coverage. For example, for sample C, the analysis system may normalize all coverage across genomic loci by the average coverage, resulting in normalized coverage. To elaborate further, for genomic locus 1, the analysis system calculates the normalized coverage for sample C by dividing the coverage at genomic locus 1 by the average coverage of sample C.

[0233] The analysis system generates a first distribution of coverage for cfDNA samples and a second distribution of coverage for WBC gDNA samples for each genomic locus.520 The analysis system can calculate various features or statistics of the distribution. For example, the analysis system can assess the normality of the distribution. The analysis system can further define the distribution based on the mean and standard deviation.

[0234] The analysis system identifies 525 highly discriminant genomic loci between cfDNA and WBC gDNA samples. The analysis system can assess the distance between the two distributions at each genomic locus. The distance can be calculated as the area under the curve for binary classification and / or as a Hodges-Lehmann estimator of the difference. The analysis system can identify highly discriminant genomic loci if the area under the curve is >0.99 or <0.01. Other thresholds, e.g., >0.98 or <0.02, >0.95 or <0.05, etc., may be used. The analysis system can further assess whether there is at least some threshold difference, e.g., >0.5, in the mean normalized coverage between the two distributions.

[0235] The analysis system determines the discriminative score for each genomic locus using a two-sample t-test.530 The two-sample t-test evaluates the discriminativeness of pairwise samples for each target. The discriminative score can be corrected using the Bonferroni correction. The discriminative score can be further based on the difference in means between the two distributions. The analysis system can further evaluate the goodness of fit for normality, i.e., whether the two distributions are sufficiently normal (Gaussian).

[0236] The analysis system determines a feature set of genomic loci based on their discriminative scores.535 The analysis system can rank genomic loci based on their discriminative scores. The analysis system can select genomic loci to form a feature set having discriminative scores above a threshold. In other embodiments, the analysis system may form a feature set by selecting the top number of genomic loci ranked according to their discriminative scores. The top number could be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, etc. The top number may be selected based on comparative validation tests regarding the predictive accuracy of each candidate set. The contamination model includes a paired distribution of coverage from cfDNA and WBC gDNA samples across the feature sets of genomic loci. For example, for 50 genomic loci in the feature set, the contamination model includes a first distribution of coverage for cfDNA samples and a second distribution of coverage for WBC gDNA samples for each of the 50 genomic loci, resulting in a total of 50 distributions of coverage for cfDNA samples and 50 distributions of coverage for WBC gDNA samples.

[0237] In one or more embodiments, the feature set of a genomic locus includes sequences annotated with respect to its genomic location, as shown in Table 1. Based on this feature set of a genomic locus, a panel of probes may target such a genomic locus.

[0238] [Table 1]

[0239] [Table 2]

[0240] [Table 3]

[0241] Figure 5B shows a flowchart illustrating WBC contamination prediction 540 using a contamination model in one or more embodiments. In one or more embodiments, the contamination model includes various distributions generated in the model training 500 of Figure 5A. For example, the contamination model includes a pairwise distribution of coverage across feature sets of genomic loci.

[0242] The analysis system obtains a test sample from the subject of study.545 The test sample may contain sequence reads sequenced from a biological sample taken from the subject of study. The test sample may be used to validate a contamination model or may be derived from a subject undergoing cancer classification. The biological sample may be a blood sample processed to isolate solid clumps containing plasma, as well as red blood cells and other microparticles in the blood, from the WBC buffy coat. The sequence reads may correspond to cfDNA fragments in the plasma. The analysis system applies a contamination model (e.g., one trained by training 500 in Figure 5A) to the test sample to evaluate whether the sample contains WBC contamination.

[0243] The analysis system determines the normalized coverage for each genomic locus in the feature set. Normalized coverage can be determined by first determining the coverage at each genomic locus in the feature set, and then normalizing the coverage by the average coverage of the test sample. Alternatively, normalization may be performed according to the sequencing depth.

[0244] The analysis system applies a contamination model to determine the contamination metric as the contribution rate of WBC gDNA that maximizes the likelihood of observing normalized coverage across the feature set, based on the coverage distribution for the feature set of genomic loci. The contamination model uses a maximum likelihood function, e.g.,

number

[0245] When the contamination metric for a test sample exceeds the contamination threshold, the analysis system determines whether the sample has WBC contamination 560. Since the contamination metric is calculated as the contribution rate of gDNA to the test sample, the contamination threshold can be an acceptable amount of gDNA contribution. This contamination threshold can be set as 0.05, 0.04, 0.03, 0.02, etc.

[0246] If it is determined that the test sample is contaminated with WBCs, the analysis system may implement corrective measures. For one, the analysis system may provide the healthcare provider with a notification indicating the detection of WBC contamination in the test sample. The notification may further recommend that the healthcare provider collect a subsequent biological sample for sequencing and analysis in the cancer classification workflow. The notification may further inform whether a particular clinical supply, a particular clinician, or a sample processing technique is more likely to result in WBC contamination. For two, the analysis system may refrain from using the test sample for further analysis in the cancer classification workflow to avoid false positive calls for cancer due to confusion with WBC-derived DNA. Refraining from using the sample from downstream analysis can involve filtering the training samples and improving the training of the cancer classification model.

[0247] IV. Cancer Classifier for Determining Cancer Cancer classification includes extracting genetic features and applying one or more models to the extracted features to determine a cancer prediction. The analysis system can aggregate the extracted features into a feature vector and then input this into a trained cancer prediction model to determine a cancer prediction based on the input feature vector. The cancer prediction can include one or more labels and / or one or more values. One label can be binary indicating the presence or absence of cancer in the test subject. Another label can be multi-class indicating one or more specific cancer types from a plurality of screened cancer types. One value can indicate the likelihood of the presence of cancer. Another value can indicate the likelihood of the absence of cancer. Yet another value can, among other things, indicate another prognosis of cancer. For example, this value can quantify the progression and / or invasion of cancer.

[0248] In one or more embodiments, the feature vector input into the cancer classifier is based on a set of pieces of information determined from the test sample.

[0249] In some embodiments, a cancer classifier may be a machine learning model that includes multiple classification parameters and a function that represents the relationship between a feature vector as input and a cancer prediction as output. The cancer prediction is obtained by inputting the feature vector into the function using the classification parameters. The machine learning model may be trained using training samples derived from individuals with known cancer diagnoses. The training samples may be divided into cohorts of various labels. For example, there may be cohorts of training samples for each type of cancer.

[0250] IV.A. Identifying highly informative fragments The analysis system can determine highly informative fragments of a sample using the methylation state vector of the sample. For each fragment in the sample, the analysis system can determine whether this fragment is highly informative using the methylation state vector corresponding to the fragment. In some embodiments, the analysis system calculates a p-value score for each methylation state vector that explains the probability of observing the methylation state vector in a healthy control group or other methylation state vectors with a lower probability. The process for calculating the p-value score is discussed further in Section IV.A. iP-value filtering below. The analysis system may determine fragments with methylation state vectors below a threshold p-value score as highly informative fragments. In some embodiments, the analysis system further labels fragments having at least a certain number of CpG sites with a certain threshold percentage of methylation or demethylation as highly methylated or low-methylated fragments, respectively. Highly methylated or low-methylated fragments may also be referred to as unusual fragments with extreme methylation (UFXM). In other embodiments, the analysis system can implement various other probabilistic models for determining highly informative fragments. Other examples of probabilistic models include mixed models and deep probabilistic models. In some embodiments, the analysis system may use any combination of the processes described below to identify highly informative fragments. Using the identified highly informative fragments, the analysis system may filter a set of methylation state vectors for a sample for use in other processes, for example, for use in training and developing a cancer classifier.

[0251] IV. AiP Value Filtering In some embodiments, the analysis system calculates a p-value score for each methylation state vector by comparing it to methylation state vectors from a fragment in a healthy control group. The p-value score can describe the probability of observing a methylation state that matches a methylation state vector in the healthy control group or another methylation state vector with a lower probability. To determine highly informative methylated DNA fragments, the analysis system can use a healthy control group that has a large proportion of fragments that are normally methylated. When this probabilistic analysis is performed to determine highly informative fragments, the determination can be weighted compared to the group of control subjects that make up the healthy control group. To ensure robustness in the healthy control group, the analysis system can select a number of healthy individuals that constitute a threshold for supplying samples containing DNA fragments. Figure 6A illustrates below how the analysis system generates a data structure for a healthy control group from which it can calculate p-value scores. Figure 6B illustrates how the p-value scores are calculated using the generated data structure.

[0252] Figure 6A is a flowchart illustrating a process 600 for generating a data structure for a healthy control group according to one embodiment. To generate a healthy control group data structure, the analysis system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. The analysis system can generate methylation state vectors for each fragment via, for example, the process 200 in Figure 2A 605.

[0253] For each fragment's methylation state vector, the analysis system can subdivide the methylation state vector into strings of CpG sites.610 In some embodiments, the analysis system subdivides the methylation state vector such that all resulting strings are less than a given length.610 For example, a methylation state vector of length 11 subdivided into strings of length 3 or less results in nine strings of length 3, ten strings of length 2, and eleven strings of length 1. In another example, a methylation state vector of length 7 subdivided into strings of length 4 or less may result in four strings of length 4, five strings of length 3, six strings of length 2, and seven strings of length 1. If the methylation state vector is less than or equal to a specified string length, the methylation state vector can be converted into a single string containing all of the CpG sites in the vector.

[0254] The analysis system aggregates strings for each possible CpG site and vector methylation state pattern by counting the number of strings present in the control group that have the specified CpG site as the first CpG site in the string and that have possible configurations of the methylation state.615 For example, for a given CpG site, considering a string length of 3, there are 2^3 or 8 possible string configurations. For each of the 8 possible string configurations for such a given CpG site, the analysis system aggregates how often each possible configuration of the methylation state vector appears in the control group.610 Continuing this example, this may include aggregating the following quantities: <M x M x+1 M x+2 >, <M x M x+1 , U x+2 >, ., For each start CpG site in the reference genome x , U x+1 , U x+2 > The analysis system generates a data structure 615 that stores aggregated counts for each possible configuration of the starting CpG site and string.

[0255] ​Setting an upper limit on string length offers several advantages. Firstly, depending on the maximum string length, the size of the data structures generated by the analysis system can increase dramatically in size. For example, a maximum string length of 4 means that to aggregate a string of length 4, all CpG sites must number at least 2^4. Increasing the maximum string length to 5 means that all CpG sites must number another 2^4 or 16 to aggregate, doubling the number of items to aggregate (and the computer memory required) compared to the previous string length. Reducing string size helps to keep the creation and performance of data structures reasonable from a computational and storage standpoint (e.g., for use for subsequent access, as described below). Secondly, a statistical consideration of limiting the maximum string length may be to avoid overfitting downstream models that use string counting. When long strings of CpG sites do not have a strong biological impact on outcomes (e.g., predicting abnormal conditions that predict the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic because it uses large amounts of data that may not be available, and therefore the model may be too sparse for it to perform properly. For example, calculating the probability of a conditioned abnormal state / cancer for the previous 100 CpG sites can be done using string counts in a data structure of length 100, ideally some of which will precisely match the previous 100 methylation states. However, if only sparse counts of strings of length 100 are available, the data may be insufficient to determine whether a given string of length 100 is abnormal in the test sample.

[0256] Figure 6B is a flowchart illustrating a process 630 for identifying highly informative methylated fragments from an individual, according to one embodiment. In process 630, the analysis system generates methylation state vectors 640 from the target cfDNA fragment, for example, via process 200 in Figure 2A. The analysis system can handle each methylation state vector as follows.

[0257] For a given methylation state vector, the analysis system enumerates all possible configurations of the methylation state vector having the same starting CpG site and the same length (i.e., set of CpG sites) in the methylation state vector.645 Since each methylation state is generally either methylated or demethylated, there may be two effectively possible states at each CpG site, and therefore the count of distinct possible configurations of the methylation state vector is such that a methylation state vector of length n has two possible configurations of the methylation state vector. n It may depend on a power of 2 to be associated with the possible configurations. For methylation state vectors that include an undeterminable state for one or more CpG sites, the analysis system enumerates the possible configurations of the methylation state vector when considering only the CpG sites that have an observed state.630

[0258] The analysis system calculates the probability of observing each possible configuration of the methylation state vector for an identified starting CpG site and methylation state vector length by accessing the data structure of a healthy control group.650 In some embodiments, calculating the probability of observing a given possible configuration models a joint probability calculation using Markov chain probabilities. The Markov model can be trained at least partially based on an assessment of the methylation state of each CpG site across the corresponding CpG sites of each fragment (e.g., a nucleic acid methylation fragment) in a healthy non-cancer cohort dataset having a corresponding set of CpG sites. For example, given a set of probabilities that determine the likelihood of observing a novel state in a sequence for each state in the sequence, a Markov model (e.g., a Hidden Markov Model or HMM) is used to determine the likelihood that a sequence of methylation states (e.g., containing "M" or "U") may be observed for a given nucleic acid methylation fragment among a set of nucleic acid methylation fragments. The set of probabilities can be obtained by training an HMM. Such training may involve calculating statistical parameters (e.g., the probability that a first state can transition to a second state (transition probability) and / or the probability that a given methylation state can be observed for each CpG site (output probability)) given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). The HMM may be trained using supervised learning (e.g., using samples where the underlying sequences and observed states are known) and / or unsupervised learning (e.g., Viterbi learning, maximum likelihood estimation, expectation maximization training and / or Baum-Welch training). In other embodiments, a calculation method other than Markov chain probabilities is used to determine the probability that each possible configuration of the methylation state vector is observed. For example, such a calculation method may include learned representations. The p-value threshold may be 0.01 to 0.10, or 0.03 to 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.

[0259] The analysis system calculates a p-value score for the methylation state vector using the probabilities calculated for each possible configuration.655 In some embodiments, this includes identifying the calculated probabilities corresponding to possible configurations that match the methylation state vector in question. Clearly, these may be possible configurations having the same set of CpG sites, or similarly, the same starting CpG sites and lengths as the methylation state vector. To generate the p-value score, the analysis system can sum the calculated probabilities of any possible configurations having probabilities less than or equal to the identified probability.

[0260] The p-value score can represent the probability of observing the methylation state vector of the fragment in a healthy control group, or other methylation state vectors with a lower probability. Thus, a low p-value score may generally correspond to a methylation state vector that is rare in healthy individuals and labels the fragment as more informative compared to the healthy control group. A high p-value score may generally be associated with a methylation state vector that is expected to be present in healthy individuals in a relative sense. If the healthy control group is non-cancerous, for example, a low p-value may indicate that the fragment has a more informative methylation pattern compared to the non-cancerous group, and therefore may indicate the presence of cancer in the test subject.

[0261] As described above, the analysis system can calculate the p-value score for each of several methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which fragments have highly informative methylation patterns, the analysis system can filter the set of methylation state vectors based on their p-value scores.665 In some embodiments, filtering is performed by comparing the p-value scores to a threshold and keeping only those fragments below the threshold. This threshold p-value score may be 0.1, 0.01, 0.001, 0.0001, or similar.

[0262] Based on exemplary results from Process 600, the analysis system can obtain a median (range) of 2,800 (1,500–12,000) fragments with highly informative methylation patterns for participants without cancer during training, and a median (range) of 3,000 (1,200–420,000) fragments with highly informative methylation patterns for participants with cancer during training. These filtered fragments with highly informative methylation patterns can be used for downstream analysis, as described in Sections IV.B and IV.C below.

[0263] In some embodiments, the analysis system uses a sliding window to determine possible configurations of the methylation state vector and calculate the p-value.660 Rather than enumerating possible configurations for the overall methylation state vector and calculating the p-value, the analysis system can enumerate possible configurations for a window of consecutive CpG sites and calculate the p-value. In this case, the window is shorter (in terms of CpG sites) than at least some fragments (otherwise the window is not useful). The window length may be static, user-determined, dynamic, or otherwise selected.

[0264] When calculating the p-value for a methylation state vector larger than the window, the window can identify a contiguous set of CpG sites within the vector, starting from a first CpG site in the vector. The analysis system can calculate the p-value score for the window containing the first CpG site. The analysis system can then "slide" the window within the vector to a second CpG site and calculate another p-value score for the second window. Thus, for a window size l and methylation vector length m, each methylation state vector can generate an m-l+1 p-value score. After the calculation of the p-value for each part of the vector is complete, the lowest p-value score from all sliding windows can be considered the overall p-value score for the methylation state vector. In other embodiments, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.

[0265] Using a sliding window can help reduce the number of possible configurations of the methylation state vector that need to be enumerated, and the number of corresponding probability calculations that would otherwise need to be performed. To give a practical example, a fragment may have more than 54 CpG sites. Instead of calculating probabilities for 2^54 (approximately 1.8 × 10^16) possible configurations to generate a single p-score, the analysis system could instead use a window of size 5 (for example), which would result in 50 p-value calculations for each of the 50 windows of the methylation state vector for that fragment. Each of the 50 calculations enumerates 2^5 (32) possible configurations of the methylation state vector, resulting in a total of 50 × 2^5 (1.6 × 10^3) ​​probability calculations. This significantly reduces the number of calculations performed, and there is no significant decrease in the accuracy of identifying highly informative fragments.

[0266] In embodiments with an undeterminable state, the analysis system can calculate a p-value score by summing the CpG sites with the undeterminable state in the methylation state vector of the fragment. The analysis system can identify all possible configurations that have consensus with all methylation states in the methylation state vector excluding the undeterminable state. The analysis system can assign probabilities to the methylation state vector as the sum of the probabilities of the identified possible configurations. For example, if the methylation states for CpG sites 1-3 are observed and there is consensus with the methylation state of the fragment at CpG sites 1-3, the analysis system can:<M1、M2、U3> and<M1、U2、U3> As the sum of the probabilities for the possible configurations of the methylation state vector,<M1、I2、U3> The probability of a methylation state vector can be calculated. This method, which sums the CpG sites that have undeterminable states, can use the calculation of the probability of up to 2^i possible configurations, which represent the number of undeterminable states in the methylation state vector. In an additional embodiment, a dynamic programming algorithm can be implemented to calculate the probability of a methylation state vector having one or more undeterminable states. Advantageously, the dynamic programming algorithm operates in linear computation time.

[0267] In some embodiments, the computational load of calculating probabilities and / or p-value scores can be further reduced by caching at least some calculations. For example, the analysis system can cache the temporary or persistent memory calculations of the probabilities (or windows thereof) for possible configurations of methylation state vectors. If other fragments have the same CpG site, by caching the probabilities of possible configurations, the p-score value can be efficiently calculated without having to recalculate the probabilities of the underlying possible configurations. Similarly, the analysis system can calculate the p-value score for each possible configuration of the methylation state vector related to a set of CpG sites from a vector (or window thereof). The analysis system can cache the p-value scores for use in determining the p-value scores of other fragments that include the same CpG site. Generally, the p-value scores of possible configurations of methylation state vectors having the same CpG site can be used to determine the p-value scores of one different possible configuration from the same set of CpG sites.

[0268] One or more nucleic acid methylation fragments can be filtered before training the region model or cancer classifier. Filtering the nucleic acid methylation fragments can include removing each nucleic acid methylation fragment that does not meet one or more selection criteria (e.g., falls below or above one or more selection criteria) from the corresponding plurality of nucleic acid methylation fragments. The one or more selection criteria can include a p-value threshold. The output p-value of each nucleic acid methylation fragment can be determined based at least in part on a comparison of the corresponding methylation pattern of each nucleic acid methylation fragment with the corresponding distribution of the methylation patterns of those nucleic acid methylation fragments in a healthy non-cancer cohort dataset having the corresponding plurality of CpG sites of each nucleic acid methylation fragment.

[0269] Filtering multiple nucleic acid methylation fragments may involve removing each nucleic acid methylation fragment that does not satisfy a p-value threshold. The filter can be applied to the methylation pattern of each nucleic acid methylation fragment using the methylation pattern observed across the first set of nucleic acid methylation fragments. Each methylation pattern of each nucleic acid methylation fragment (e.g., fragment 1, ..., fragment N) may include one or more corresponding methylation sites (e.g., CpG sites) identified by a methylation site identifier, and a corresponding methylation pattern represented as a sequence of 1' and 0', where each "1" represents a methylated CpG site in one or more CpG sites, and each "0" represents an unmethylated CpG site in one or more CpG sites. Using the methylation pattern observed across the first set of nucleic acid methylation fragments, a methylation state distribution of CpG site states (e.g., CpG site A, CpG site B, ..., CpG site ZZZ) collectively represented by the first set of nucleic acid methylation fragments can be constructed. Further details relating to the processing of nucleic acid methylated fragments are disclosed in U.S. Provisional Patent Application No. 17 / 191,914, filed on March 4, 2021, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," which is incorporated herein by reference.

[0270] Each nucleic acid methylation fragment may not meet one or more selection criteria if it has an informative methylation score below an informative methylation score threshold. In this situation, the informative methylation score can be determined by a mixed model. For example, a mixed model can detect informative methylation patterns in nucleic acid methylation fragments by determining the likelihood of the methylation state vector (e.g., methylation pattern) for each nucleic acid methylation fragment based on the number of possible methylation state vectors of the same length for the same corresponding genomic location. This can be done by generating multiple possible methylation states for a vector of a specified length at each genomic location in the reference genome. Using the multiple possible methylation states, the total number of possible methylation states, followed by the probability of each predicted methylation state at the genomic location, can be determined. The likelihood of the sample nucleic acid methylation fragment corresponding to the genomic location in the reference genome can then be determined by matching the sample nucleic acid methylation fragment to the predicted (e.g., possible) methylation states and by retrieving the calculated probability of the predicted methylation states. The informative methylation score can then be calculated based on the probability of the sample nucleic acid methylation fragment.

[0271] Each nucleic acid methylation fragment may fail to meet one or more selection criteria if it contains fewer than a threshold number of residues. The threshold number of residues may be 10–50, 50–100, 100–150, or greater than 150. The threshold number of residues may be a fixed value of 20–90. Each nucleic acid methylation fragment may fail to meet one or more selection criteria if it contains fewer than a threshold number of CpG sites. The threshold number of CpG sites may be 4, 5, 6, 7, 8, 9, or 10. Each nucleic acid methylation fragment may fail to meet one or more selection criteria if its genome start and end locations indicate that it represents fewer than a threshold number of nucleotides in the human genome reference sequence.

[0272] Filtering can remove nucleic acid methylation fragments in a group of corresponding nucleic acid methylation fragments that have the same methylation pattern as another nucleic acid methylation fragment in the group of corresponding nucleic acid methylation fragments, as well as nucleic acid methylation fragments in the group of corresponding nucleic acid methylation fragments that have the same corresponding genome start and end locations. This filtering step can remove redundant fragments that are exact duplicates, possibly including PCR repeats. Filtering can remove nucleic acid methylation fragments that have the same corresponding genome start and end locations as another nucleic acid methylation fragment in the group of corresponding nucleic acid methylation fragments, and that have fewer than a threshold number of different methylation states. The threshold number of different methylation states used to retain nucleic acid methylation fragments may be 1, 2, 3, 4, 5, or greater than 5. For example, a first nucleic acid methylation fragment (e.g., aligned to a reference genome) is retained that has the same corresponding genome start and end locations as a second nucleic acid methylation fragment, but has at least one, at least two, at least three, at least four, or at least five different methylation states at each CpG site. As another example, a first nucleic acid methylation fragment is also retained, which has the same methylation state vector (e.g., methylation pattern) but has different corresponding genome start and end positions than the second nucleic acid methylation fragment.

[0273] Filtering can remove assay artifacts in multiple nucleic acid methylation fragments. Removal of assay artifacts may include removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences that did not undergo conversion during bisulfite conversion. Filtering can remove contaminants (e.g., resulting from sequencing, nucleic acid isolation, and / or sample preparation).

[0274] Filtering can remove subsets of methylation fragments from multiple methylation fragments based on cross-information filtering of each methylation fragment for cancer status across multiple training subjects. For example, cross-information can provide a measure of interdependence between two conditions of interest sampled simultaneously. Cross-information can be determined by selecting an independent set of CpG sites (e.g., all or part of nucleic acid methylation fragments) from one or more datasets, and by comparing the probabilities of methylation states for sets of CpG sites between two sample groups (e.g., genotype datasets, biological samples and / or subsets and / or groups of subjects). The cross-information score can represent the probability of a methylation pattern for a first condition relative to a second condition in each region within each frame of a sliding window, and thus indicate the discriminative power of each region. The cross-information score can be similarly calculated for each region within each frame of the sliding window as it progresses across the selected set of CpG sites and / or selected genomic regions. Further details regarding mutual information filtering are disclosed in U.S. Patent Application No. 17 / 119,606, filed on 11 December 2020, entitled "Cancer Classification Using Patch Convolutional Neural Networks," which is incorporated herein by reference in its entirety.

[0275] IV.A.ii. Highly methylated and lowly methylated fragments In some embodiments, the analysis system identifies low-methylated or high-methylated fragments from a filtered set as highly informative fragments.670 The analysis system identifies high-methylated fragments having a threshold number of CpG sites and a threshold percentage of methylated CpG sites. The analysis system identifies low-methylated fragments having a threshold number of CpG sites and a threshold percentage of unmethylated CpG sites. Exemplary thresholds for fragment (or CpG site) length include values ​​exceeding 3, 4, 5, 6, 7, 8, 9, 10, etc. Exemplary threshold percentages for methylation or unmethylation include values ​​exceeding 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%.

[0276] IV.B. Training of the cancer classifier Figure 7A is a flowchart illustrating a process 700 for training a cancer classifier according to one embodiment. The analysis system obtains multiple training samples 710, each having a set of highly informative fragments and a cancer type label. The multiple training samples may include any combination of samples from healthy individuals with a general label of “non-cancerous,” and samples from subjects with a general label of “cancer” or a specific label (e.g., “breast cancer,” “lung cancer,” etc.). Training samples from subjects for one cancer type may be called a cohort for that cancer type or a cancer type cohort.

[0277] The analysis system determines a feature vector for each training sample based on a set of highly informative fragments of the training sample.720 The analysis system can calculate the informality score of each CpG site in an initial set of CpG sites. The initial set of CpG sites may be all CpG sites in the human genome or a portion thereof, and this is 10 4 , 10 5 , 10 6 , 10 7 , 10 8The degree of informationality may vary. In one embodiment, the analysis system defines an informationality score for a feature vector using a binary score based on whether there is an informational fragment within a set of informational fragments that encompass the CpG site. In another embodiment, the analysis system defines an informationality score based on the count of informational fragments that overlap with the CpG site. For example, the analysis system may use ternary scoring, assigning a first score for the absence of informational fragments, a second score for the presence of several informational fragments, and a third score for the presence of more than a few informational fragments. For example, the analysis system counts five informational fragments in the sample that overlap with the CpG site and calculates an informationality score based on the count of five.

[0278] Once all information scores have been determined for the training sample, the analysis system can determine a feature vector as a vector of elements containing one of the information scores associated with one of the CpG regions in the initial set for each element. The analysis system can normalize the information scores of the feature vector based on the sample coverage, where coverage may refer to the median or mean sequence depth across all CpG regions covered by the initial set of CpG regions used by the classifier, or based on a set of highly informative fragments of a given training sample.

[0279] As an example, see Figure 7B, which shows the matrix of the training feature vector 722. In this example, the analysis system identifies CpG sites [K]726 that should be considered in generating the feature vector for the cancer classifier. The analysis system selects the training sample [N]724. The analysis system determines a first informability score 728 for a first arbitrary CpG site [k1] to be used in the feature vector for the training sample [n1]. The analysis system checks each informative fragment in the set of informative fragments. If the analysis system identifies at least one informative fragment containing the first CpG site, the analysis system determines the first informability score 728 for the first CpG site to be 1, as shown in Figure 7B. If a second arbitrary CpG site [k2] is considered, the analysis system similarly checks the set of informative fragments for at least one containing the second CpG site [k2]. If the analysis system cannot find any such highly informative fragment containing the second CpG region, the analysis system determines a second informality score of 729 for the second CpG region [k2] to be 0, as shown in Figure 7B. Once the analysis system has determined all informality scores for the initial set of CpG regions, the analysis system determines a feature vector for the first training sample [n1] that includes an informality score of 728 of 1 for the first CpG region [k1] and a second informality score of 729 of 0 for the second CpG region [k2], and subsequent informality scores, thereby forming the feature vector [1, 0, ...].

[0280] Additional approaches to sample characterization are described in U.S. Patent Application No. 15 / 931,022 entitled "Model-Based Featurization and Classification," U.S. Patent Application No. 16 / 579,805 entitled "Mixture Model for Targeted Sequencing," U.S. Patent Application No. 16 / 352,602 entitled "Anomalous Fragment Detection and Classification," and U.S. Patent Application No. 16 / 723,716 entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples," all of which are incorporated in their entirety by reference.

[0281] The analysis system may further restrict the CpG sites considered for use in the cancer classifier. For each CpG site in the initial set of CpG sites, the analysis system calculates the information gain based on the feature vector of the training sample. From step 720, each training sample has a feature vector that may contain information scores for all CpG sites in the initial set of CpG sites, which may include up to all CpG sites in the human genome. However, some CpG sites in the initial set of CpG sites may not be as informational as others in distinguishing cancer types, or they may overlap with other CpG sites.

[0282] In one embodiment, the analysis system calculates the information gain for each cancer type and each CpG site in the initial set to determine whether to include that CpG site in the classifier.730 The information gain is calculated to train a sample having a given cancer type compared to all other samples. For example, two random variables “informative fragment” ("IF") and “cancer type” ("CT") are used. In one embodiment, IF is a binary variable indicating whether there is an informative fragment in a given sample that overlaps with a given CpG site, as determined for the informality score / feature vector. CT is a random variable indicating whether the cancer is of a particular type. The analysis system calculates the mutual information for CT given AF; that is, how many bits of information about the cancer type are obtained if it is known whether there is an informative fragment that overlaps with a particular CpG site. In practice, for a first cancer type, the analysis system calculates the paired mutual information gain for each cancer type and sums the mutual information gains across all other cancer types.

[0283] For a given cancer type, the analysis system can use this information to rank CpG sites based on their cancer specificity. This procedure can be repeated for all cancer types considered. If certain regions are generally highly informative and methylated in a given cancer training sample, but not methylated in training samples of other cancer types or healthy training samples, CpG sites that overlap with those highly informative fragments can provide high information about the given cancer type. Ranked CpG sites for each cancer type can be greedily added (selected) to a selected set of CpG sites based on their rank for use in the cancer classifier.

[0284] In additional embodiments, the analysis system may consider other selection criteria for selecting highly informative CpG sites to be used in the cancer classifier. One selection criterion may be that the selected CpG site exceeds a threshold interval from other selected CpG sites. For example, the selected CpG site must exceed a threshold number of base pairs (e.g., 100 base pairs) away from any other selected CpG site, and as a result, CpG sites within the criterion interval are both not selected for consideration in the cancer classifier.

[0285] In one embodiment, according to a selected set of CpG sites from the initial set, the analysis system may modify the feature vectors of the training samples as needed.750 For example, the analysis system may truncate the feature vectors to remove informality scores corresponding to CpG sites that are not in the selected set of CpG sites.

[0286] Using the feature vectors of the training samples, the analysis system can train a cancer classifier in one of several ways. The feature vectors may correspond to an initial set of CpG sites from step 720 or a selected set of CpG sites from step 750. In one embodiment, the analysis system trains a binary cancer classifier to distinguish between cancer and non-cancerous samples based on the feature vectors of the training samples 760. In this way, the analysis system uses training samples that include both non-cancerous samples from healthy individuals and cancerous samples from subjects. Each training sample may have one of two labels, "cancer" or "non-cancerous". In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.

[0287] In another embodiment, the analysis system trains a multi-class cancer classifier to distinguish between multiple cancer types (also called tissue of origin (TOO) labels). Cancer types may include one or more cancers and may include non-cancer types (including any further other diseases or genetic disorders, etc.). To do this, the analysis system may use cohorts of cancer types, and may or may not include cohorts of non-cancer types. In this multi-cancer embodiment, the cancer classifier is trained to determine cancer predictions (or more specifically, TOO predictions) that include prediction values ​​for each of the classified cancer types. The prediction values ​​may correspond to the likelihood that a given training sample (and test sample during inference) has each of the cancer types. In one embodiment, the prediction values ​​are scored between 0 and 100, with the cumulative prediction value being 100. For example, the cancer classifier returns cancer predictions that include prediction values ​​for breast cancer, lung cancer, and non-cancer. For example, the classifier might return a cancer prediction that the test sample has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of being non-cancerous. The analysis system further evaluates the prediction values ​​to generate a prediction of the presence of one or more cancers in the sample, which may also be referred to as a TOO prediction, indicating one or more TOO labels, such as a first TOO label with the highest prediction value, a second TOO label with the second highest prediction value, and so on. Continuing the above example, and given these percentages, in this example the system might determine that the sample has breast cancer, considering that breast cancer has the highest likelihood.

[0288] In both embodiments, the analysis system trains the cancer classifier by inputting a set of training samples containing the feature vectors being trained by the classifier function into the cancer classifier and adjusting the classification parameters so that the feature vectors being trained by the classifier function accurately relate to their corresponding labels. The analysis system can group the training samples into one or more sets of training samples for iterative batch training of the cancer classifier. After inputting all sets of training samples containing those training feature vectors and adjusting the classification parameters, the cancer classifier can be sufficiently trained to label test samples according to the feature vectors within a certain margin of error. The analysis system can train the cancer classifier according to one of several methods. As an example, a binary cancer classifier may be an L2 regularized logistic regression classifier trained using a log-loss function. As another example, a multi-cancer classifier may be a multinomial logistic regression. In practice, any type of cancer classifier can be trained using other techniques. These techniques are numerous and include the potential use of machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autoencoder models, and multilayer neural networks.

[0289] The classifier may include logistic regression algorithms, neural network algorithms, support vector machine algorithms, naive Bayes algorithms, nearest neighbor algorithms, boosted tree algorithms, random forest algorithms, decision tree algorithms, multinomial logistic regression algorithms, linear models, or linear regression algorithms.

[0290] Development of IVC cancer classifiers During the use of the cancer classifier, the analysis system can obtain test samples from subjects of unknown cancer types. The analysis system can process test samples consisting of DNA molecules in any combination of processes 200 and 630 to obtain a set of highly informative fragments. The analysis system can determine a test feature vector for use by the cancer classifier according to a similar principle discussed in process 700. The analysis system can calculate an informality score for each CpG site within multiple CpG sites when used by the cancer classifier. For example, the cancer classifier receives an input feature vector containing informality scores for 1,000 selected CpG sites. Thus, the analysis system can determine a test feature vector containing informality scores for 1,000 selected CpG sites based on the set of highly informative fragments. The analysis system can calculate the informality scores in the same manner as for the training samples. In some embodiments, the analysis system defines the informality score as a binary score based on whether highly methylated or low methylated fragments are present in the set of highly informative fragments containing the CpG sites.

[0291] Next, the analysis system can input the test feature vector into a cancer classifier. The cancer classifier function can then generate cancer predictions based on the classification parameters and test feature vectors trained in process 700. In the first method, the cancer prediction can be binary and selected from a group consisting of "cancer" or non-cancer, while in the second method, the cancer prediction is selected from a group consisting of many cancer types and "non-cancer". In an additional embodiment, the cancer prediction has a prediction value for each of the many cancer types. In addition, the analysis system can determine that the test sample is most likely to be one of the cancer types. Following the above embodiment, using cancer predictions for the test sample with a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of non-cancer, the analysis system can determine that the test sample is most likely to have breast cancer. In another example, if the cancer prediction is binary with a 60% likelihood of non-cancer and a 40% likelihood of cancer, the analysis system determines that the test sample is most likely to not have cancer. In additional embodiments, to refer to the test subject as having that type of cancer, the cancer prediction with the highest likelihood can still be compared to a threshold (e.g., 40%, 50%, 60%, 70%). If the cancer prediction with the highest likelihood does not exceed that threshold, the analysis system may return an indeterminate result.

[0292] In additional embodiments, the analysis system chains a cancer classifier trained in step 760 of process 700 with another cancer classifier trained in step 770 or process 700. The analysis system can input the test feature vector to a cancer classifier trained as a binary classifier in step 760 of process 700. The analysis system can receive the output of the cancer prediction. The cancer prediction may be binary, indicating whether the subject is likely to have cancer or not. In other embodiments, the cancer prediction includes prediction values ​​representing the likelihood of cancer and the likelihood of not having cancer. For example, the cancer prediction may have an 85% cancer prediction value and a 15% non-cancer prediction value. The analysis system may determine that the subject is likely to have cancer. If the analysis system determines that the subject is likely to have cancer, the analysis system can input the test feature vector to a multiclass cancer classifier trained to distinguish between different types of cancer. The multiclass cancer classifier can receive the test feature vector and return a cancer prediction for one of the multiple types of cancer. For example, a multiclass cancer classifier provides a cancer prediction that identifies the subject of study as most likely to have ovarian cancer. In another embodiment, the multiclass cancer classifier provides a prediction value for each of several cancer types. For example, the cancer prediction may include a 40% prediction for breast cancer, a 15% prediction for colorectal cancer, and a 45% prediction for liver cancer.

[0293] According to a generalized embodiment of binary cancer classification, an analysis system can determine a cancer score for a test sample based on the sample's sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analysis system can compare the cancer score for the test sample to a binary threshold cutoff to predict whether the test sample is likely to have cancer. The binary threshold cutoff can be adjusted using TOO thresholding based on one or more TOO subtype classes. The analysis system can further generate feature vectors for the test sample for use in a multi-class cancer classifier to determine a cancer prediction indicating one or more possible cancer types.

[0294] A classifier can be used to determine the disease status of a test subject, for example, a subject whose disease status is unknown. This method may include observing the test's genomic data construct (e.g., test data at a single time point), which, in electronic form, contains values ​​for each genomic feature among multiple genomic features of corresponding nucleic acid fragments in a biological sample obtained from the test subject. The method may then include applying the test's genomic data construct to a test classifier to determine the disease status of the test subject. The test subject may not have been previously diagnosed with a disease.

[0295] The classifier may be a temporal classifier that uses at least (i) a first test genome data construct generated from a first biological sample obtained from a test subject at a first time point, and (ii) a second test genome data construct generated from a second biological sample obtained from a test subject at a second time point.

[0296] Using a trained classifier, for example, the disease status of a test subject, whose condition is unknown, can be determined. In this case, the method includes obtaining a test time-series dataset in electronic format for the test subject, which includes a corresponding test genotype data construct for each of several time points, containing values ​​for multiple genotype features of multiple corresponding nucleic acid fragments in the corresponding biological sample obtained from the test subject at each time point, and for each pair of consecutive time points, including an indication of the length of time between the consecutive time points in the pair. The method may then include applying the test genotype data construct to a test classifier to determine the disease status of the test subject. The test subject may not have been previously diagnosed with a disease.

[0297] V.Application In some embodiments, the methods, analytical systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, determine the presence of cancer or monitor minimum residual disease (MRD), or any combination thereof. For example, as described herein, the classifier can be used to generate a probability score (e.g., 0 to 100) that describes the likelihood that a test feature vector originates from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether or not a subject has cancer. In other embodiments, the likelihood or probability score can be evaluated at a number of different time points (e.g., before and after treatment) to monitor disease progression or the effectiveness of treatment (e.g., treatment effect). In yet another embodiment, the likelihood or probability score can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, evaluation of treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe appropriate treatment. In additional embodiments, the methods, analytical systems, and / or classifiers can be implemented to detect sources of contamination in sample handling and analytical workflows. Upon detecting contamination and / or its source, the analysis system can assist in implementing corrective actions to mitigate the contamination and its adverse effects (e.g., distortion of results, bias training of classifiers, etc.).

[0298] VAWBC contamination In some embodiments, the method and / or classifier of the present invention is used to detect WBC contamination, for example, as described by the various methodologies in Section III.

[0299] According to one or more embodiments of the first methodology for detecting WBC contamination, Figures 9A to 14 illustrate exemplary results.

[0300] Figure 9A shows a volcano plot of coverage ratios for an initial set of genomic loci under consideration. The coverage ratio is calculated by dividing the coverage between WBC samples and cfDNA samples. The x-axis plots the logarithmic scaling of the coverage ratio. For example, a logarithmic scaling with a base-2 logarithmic scale results in a value of 1.5, which corresponds to 2^1.5, approximately equal to a coefficient of 2.82. The y-axis plots this as the significance of the genomic loci, i.e., the logarithm of the p-value. A larger value on the y-axis indicates higher statistical significance. Here, genomic loci were selected within the upper right square (e.g., shown as triangles) as those with coverage ratios exceeding a threshold and high statistical significance.

[0301] Figure 9B shows a visual comparison of differential coverage at selected genomic loci between cfDNA and WBC samples. Genomic loci are listed vertically, e.g., along the y-axis, and samples are arranged horizontally, e.g., along the x-axis. The legend below indicates that darker areas represent lower normalized coverage and brighter areas represent higher normalized coverage. Notably, horizontal light bands are consistently present in the WBC samples, while corresponding light bands are absent in the cfDNA samples. This demonstrates the discriminative power of the feature sets at genomic loci.

[0302] Figure 10 shows a single genomic locus in mitochondrial DNA that exhibits a striking difference in representation between cfDNA and WBC samples. The distribution of fragments matching chrM (mitochondrial DNA) and chr10 (control) in whole-genome sequencing data for cfDNA and shear leukocyte gDNA from CCGA1 samples. The left graph shows that the proportion of cfDNA matching chrM (red and left) is significantly lower (approximately 17 times) than that observed in WBC gDNA (blue and right) (note the log10 scale on the x-axis). The right graph, as a control, shows that the difference between the proportion of chr10-matching fragments between cfDNA and WBC DNA is even more similar (within 5%). The low proportion of cfDNA fragments matching chrM is likely due to the lack of nucleosomes in mitochondria, which protect this DNA from in vivo degradation after efflux from the cell, unlike most cfDNA derived from nuclear DNA that is typically surrounded by nucleosomes. Such differences in coverage in chrM can be used to monitor WBC contamination in clinical samples, as gDNA eluted from WBCs into plasma collected in EDTA or STRECK tubes is not as rapidly degraded as in peripheral blood. As a result, chrM, for example, 1e above -4 An increase in the proportion of fragments matching the criteria indicates WBC contamination in the sample.

[0303] Figure 11 shows a comparison of significant genomic loci in feature sets of genomic loci compared to randomly selected candidate sets. Figure 11 shows a bar graph quantifying how many candidate sets (out of 100 candidate sets) have varying numbers of significant genomic loci. The threshold for a locus to be considered significant is whether the ratio of genomic locus coverage by WBC samples to genomic locus coverage by cfDNA samples is greater than 2. Nearly 70 candidate sets had 0 significant genomic loci and subsequently had little to no predictive ability in assessing the presence of WBC contamination. Approximately 22 candidate sets had one significant genomic locus, and some candidate sets had more than one significant genomic locus but fewer than 5. Feature sets (e.g., identified by feature identification 300 in Figure 3A) obtained at least 15 significant genomic loci. This bar graph shows the improvement of feature sets compared to randomly selected markers or genomic loci.

[0304] Figure 12 shows an exemplary pair of WBC and cfDNA samples from the same subject, evaluated across feature sets at genomic loci (e.g., identified by feature identification 300 in Figure 3A). Notably, no single target or genomic locus can simply predict WBC contamination. Rather, across various genomic loci, several pairs of samples exhibit similar coverage in both WBC and cfDNA samples, providing no discriminative power. Nevertheless, the aggregation of various genomic loci in the feature sets is linked to obtain the predictive power of coverage-based WBC contamination methodologies.

[0305] In the first set of validation tests, samples enriched with longer fragments (up to the top 1% of sequence reads) were tested against a set of genomic locus features for coverage-based WBC contamination against a randomly selected candidate set of genomic loci. The feature set (as identified in Feature Identification 300 in Figure 3A) identified 3 outlier samples out of 100. The results for the candidate sets are summarized in Table 2 below.

[0306] [Table 4]

[0307] The majority of the candidate sets did not predict any outlier samples, i.e., no samples were predicted to be included in the WBC, with just under half of the candidate sets predicting one outlier sample. Approximately 3.5% of the candidate sets predicted two outlier samples, and only one out of 500 candidates (0.2%) predicted three outlier samples.

[0308] In the second set of validation tests, samples enriched with longer fragments (up to the top 1% of sequence reads) were tested against a set of genomic locus features for coverage-based WBC contamination against a randomly selected candidate set of genomic loci. The feature set (as identified in Feature Identification 300 in Figure 3A) identified 7 outlier samples out of 62. The results for the candidate sets are summarized in Table 3 below.

[0309] [Table 5]

[0310] The majority of the candidate sets (60%) did not predict any outlier samples, i.e., no samples were predicted to be contaminated with WBCs. Approximately one-third (36%) of the candidate sets predicted one outlier sample, and 4% predicted two outlier samples. No candidate set predicted more than two outlier samples, and the feature set identified seven outlier samples.

[0311] Figure 13 shows samples identified as containing WBCs by coverage-based WBC contamination detection. Along the x-axis, samples are plotted based on high molecular weight. Along the y-axis, samples are plotted based on z-scores, as determined by inference 370 in Figure 3C. White triangles indicate outliers, i.e., samples most confidently determined to be containing WBCs, for the genomic locus feature set (e.g., as identified according to feature identification 300 in Figure 3A). Black triangles indicate samples identified as outliers by the genomic locus feature set, but most confidently determined to be outliers by the random candidate set. Many of the white triangles were not identified as containing WBCs by the other random candidate sets.

[0312] Figure 14 shows samples in which the longest fragments (top 1%) were identified as containing WBCs by coverage-based WBC contamination detection. Along the x-axis, samples are plotted based on the total length of the fragments in the sample. Along the y-axis, samples are plotted based on z-scores, as determined by inference 370 in Figure 3C. White triangles indicate outliers for the feature set of genomic loci, i.e., samples that were most confidently determined to be WBC contaminated (e.g., as identified according to feature identification 300 in Figure 3A). Importantly, fragment length is not a reliable indicator of WBC contamination. Many samples with a fragment length above 360 ​​in the average top 1% were not WBC contaminated. In addition, some samples with a fragment length below the 360 ​​mark were WBC contaminated. Using a fragment length cutoff is not accurate in WBC contamination detection.

[0313] According to one or more exemplary embodiments of the second methodology for detecting WBC contamination, Figures 15 to 19 illustrate exemplary results.

[0314] Figure 15 shows the contribution rates of various tissue types for three different sample types: cfDNA sample, purified plasma sample, and serum sample containing only WBC DNA, according to an exemplary embodiment. The tissue types include epithelial cell type, hepatocyte type, lymphocyte cell type, bone marrow circulating tissue type, bone marrow noncirculating tissue type, and vascular endothelial cell type. Notably, the contribution rate of plasma is not similar to that of serum. This is a proof of concept for an embodiment of the second methodology.

[0315] Figure 16 shows the information gain across the set of genomic loci evaluated for inclusion as part of the deconvolution model (as trained in Figure 4A). The x-axis plots the mutual information obtained during training of the deconvolution model, and the y-axis plots the prediction error (mean squared error). Each graph represents the mutual information obtained for a single genomic locus or deconvolution marker.

[0316] Figure 17 shows box plots illustrating the predictive power of a methylation-based WBC contamination detection methodology. The y-axis of the left graph plots the Mahalanobis distance, i.e., the distance from the mean proportion of tissue types for the null distribution of cfDNA samples from the non-cancer population. The y-axis of the right graph plots the p-value as a contamination metric for calling WBC contamination in a sample. The contamination threshold is set to 0.05. The green box plot labeled CCGA1 cfDNA (left side of each graph) refers to cfDNA samples from the CCGA1 dataset. The Mahalanobis distance (used to calculate the contamination metric) is relatively low among the cfDNA samples, and correspondingly the contamination metric is generally very high, with only a few selected outliers below the contamination threshold. These samples may have been contaminated with WBC without any positive detection. The orange box plot labeled Columbia Serum (right side of each graph) refers to WBC gDNA samples. In contrast to cfDNA samples, WBC samples exhibit larger Mahalanobis distances and generally lower contamination metrics as p-values. Notably, the majority of WBC gDNA samples would be considered WBC contamination, with p-values ​​below a contamination threshold of 0.05.

[0317] Figure 18 shows titration samples mixed with cfDNA and WBC serum according to an exemplary embodiment. Each graph has the same axis, with the x-axis representing tissue type (epithelial cell type, hepatocyte type, lymphocyte type, circulating bone marrow tissue type, non-circulating bone marrow tissue type, and vascular endothelial cell type) and the y-axis representing the contribution rate of each tissue type. The baseline experiment (upper left graph without WBC serum addition) shows, on average, a 50% contribution rate from non-circulating bone marrow tissue types, approximately 37.5% from circulating bone marrow tissue types, approximately 5% from lymphocyte type, and less than 5% for the remaining tissue types. As more serum was titrated into the sample, the contribution rate of non-circulating bone marrow tissue types increased, accompanied by a decrease in the contribution rates of non-circulating bone marrow tissue types and lymphocyte type. These experiments validate the predictions in Figure 15.

[0318] Figure 19 shows titration samples with Mahalanobis distances calculated according to the second methodology. Each graph corresponds to a single subject with exemplary titration of WBC serum into cfDNA plasma. The x-axis represents the titration level of WBC serum, and the y-axis represents the Mahalanobis distance. The majority of samples show an increase in Mahalanobis distance, which indicates a further deviation from the mean tissue type proportion in the null distribution of cfDNA samples from non-cancer subjects. This supports the predictive power of utilizing Mahalanobis distance in contamination metric calculations.

[0319] According to one or more exemplary embodiments of the third methodology for detecting WBC contamination, Figure 20 illustrates exemplary results.

[0320] Figure 20 shows titration samples quantified for WBC contamination using the third methodology. Each point represents a titration sample containing a set percentage of WBC serum DNA titrated into cfDNA plasma. The x-axis plots the percentage of WBC serum in each sample, while the y-axis plots the predicted gDNA contamination based on a model trained according to the third methodology described in Figures 5A and 5B. Notably, the model predicted the increasing WBC gDNA contamination with increasing titration class progressively accurately. For example, the spread of WBC gDNA contamination in the 0.10 titration class was 0.00 to approximately 0.16, while the spread of WBC gDNA contamination in the 0.30 titration class was approximately 0.13 to approximately 3.2. The spread also showed considerable consistency across titration classes.

[0321] Early detection of VB cancer In some embodiments, the methods and / or classifiers of the present invention are used to detect the presence or absence of cancer in subjects suspected of having cancer. For example, a classifier (e.g., described above in Section IV and illustrated in Section V) can be used to determine a cancer prediction that explains the likelihood that the test feature vector originates from a subject having cancer.

[0322] In one embodiment, cancer prediction is the likelihood (e.g., scored between 0 and 100) of whether a test sample has cancer (i.e., a binary classification). Thus, the analysis system can determine a threshold for determining whether a test subject has cancer. For example, a cancer prediction of 60 or higher may indicate that the subject has cancer. In yet another embodiment, cancer predictions of 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, or 95 or higher may indicate that the subject has cancer. In yet another embodiment, cancer prediction can indicate the severity of the disease. For example, a cancer prediction of 80 may indicate a more severe form or a later stage of cancer compared to a cancer prediction of less than 80 (e.g., a probability score of 70). Similarly, an increase in cancer prediction over time (e.g., determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate successful treatment.

[0323] In another embodiment, the cancer prediction includes many prediction values, and each of the multiple cancer types that are classified (i.e., multi-class classification) has a prediction value (e.g., scored between 0 and 100). The prediction values ​​may correspond to the likelihood that a given training sample (and training sample during inference) has each of the cancer types. The analysis system can identify the cancer type with the highest prediction value and indicate that the subject of study is likely to have that cancer type. In another embodiment, the analysis system further compares the highest prediction value with thresholds (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to determine that the subject of study is likely to have the cancer type. In another embodiment, the prediction values ​​may also indicate the severity of the disease. For example, a prediction value above 80 may indicate a more severe form or a later stage of cancer compared to a prediction value of 60. Similarly, an increase in predicted values ​​over time (determined, for example, by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in predicted values ​​over time may indicate treatment success.

[0324] According to aspects of the present invention, the methods and systems of the present invention can be trained to detect or classify multiple cancer signs. For example, the methods, systems and classifiers of the present invention can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different cancer types.

[0325] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include carcinomas, lymphomas, blastomas, sarcomas, and leukemias or lymphoid malignancies. More specific examples of such cancers include squamous cell carcinoma (e.g., epithelial squamous cell carcinoma), skin cancer, malignant melanoma, lung cancer (including small cell lung cancer, non-small cell lung cancer (NSCLC), lung adenocarcinoma, and lung squamous cell carcinoma), peritoneal cancer, gastric cancer, gastrointestinal cancer, pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high-grade serous ovarian cancer), and liver cancer (e.g., hepatocellular carcinoma). Examples of cancers that fall under this category include, but are not limited to, carcinoma (HCC), hepatoma, liver cancer, bladder cancer (e.g., urothelial carcinoma), testicular cancer (germ cell tumor), breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer (e.g., renal cell carcinoma, nephroblastoma (Wilms' tumor)), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal carcinoma (NPC). Additional examples of cancers include, but are not limited to, retinoblastoma, theca cell tumor, arinoblastoma, hematological malignancies (including, but not limited to, non-Hodgkin's lymphoma (NHL), multiple myeloma, and acute hematological malignancies), endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.

[0326] In some embodiments, the cancer is one or more of the following: anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, stomach cancer, head and neck cancer, hepatobiliary tract cancer, leukemia, lung cancer, lymphoma, malignant melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, uterine cancer, or any combination thereof.

[0327] In some embodiments, one or more cancers may be anorectal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, pancreatic cancer, and "high-signal" cancers (defined as cancers with a 5-year cancer-specific mortality rate greater than 50%) such as lymphoma and multiple myeloma. High-signal cancers tend to be more invasive and typically have above-average cell-free nucleic acid concentrations in test samples obtained from patients.

[0328] VC cancer and treatment monitoring In some embodiments, cancer prediction may be evaluated at multiple different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., treatment efficacy). For example, the present invention includes a method comprising obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point and determining a first cancer prediction therefrom (as described herein), and obtaining a second test sample (e.g., a second plasma cfDNA sample) from a cancer patient at a second time point and determining a second cancer prediction therefrom (as described herein). The classification may further quantify tumor volume to evaluate changes over time.

[0329] In a particular embodiment, the first time point is before cancer treatment (e.g., before surgical resection or therapeutic intervention), and the second time point is after cancer treatment (e.g., after surgical resection or therapeutic intervention), and the classifier is used to monitor the effectiveness of the treatment. For example, if the second cancer prediction decreases compared to the first cancer prediction, the treatment is considered successful. However, if the second cancer prediction increases compared to the first cancer prediction, the treatment is considered unsuccessful. In another embodiment, both the first and second time points are before cancer treatment (e.g., before surgical resection or therapeutic intervention). In yet another embodiment, both the first and second time points are after cancer treatment (e.g., after surgical resection or therapeutic intervention). In further embodiments, cfDNA samples may be obtained from cancer patients at a first and second time point and analyzed, for example, to monitor cancer progression, to determine whether the cancer is in remission (e.g., after treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effectiveness of treatment (e.g., therapeutic agents).

[0330] Those skilled in the art will readily understand that test samples can be obtained from cancer patients over a series of arbitrary point in time and analyzed according to the method of the present invention to monitor the patient's cancer status. In some embodiments, the first and second point in time are amounts of time ranging from about 15 minutes to a maximum of about 30 years, for example, about 30 minutes, for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or about 24 hours, for example, about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 50 days, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or for example, about 1, 1.5, 2, 2. It can be divided into 5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or approximately 30 years. In other embodiments, test samples can be obtained from patients at least once every five months, at least once every six months, at least once every one year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0331] VD treatment In yet another embodiment, cancer prediction can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, evaluation of treatment effectiveness, etc.). For example, in one embodiment, if cancer prediction (e.g., cancer or a specific type of cancer) exceeds a threshold, a physician may prescribe appropriate treatment (e.g., surgical resection, radiotherapy, chemotherapy, and / or immunotherapy).

[0332] A classifier (as described herein) can be used to determine a cancer prediction that a sample feature vector is from a subject with cancer. In one embodiment, if the cancer prediction exceeds a threshold, an appropriate treatment (e.g., surgical resection or therapeutic drug) is prescribed. For example, in one embodiment, if the cancer prediction is 60 or higher, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, one or more appropriate treatments are prescribed. In yet another embodiment, the cancer prediction may indicate the severity of the disease. An appropriate treatment matching the severity of the disease may then be prescribed.

[0333] In some embodiments, the treatment is one or more cancer agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapies, differentiation induction therapies, hormone therapies, and immunotherapies. For example, the treatment may be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based drugs, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapies selected from the group consisting of signaling inhibitors (e.g., tyrosine kinases and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoin receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody-drug conjugates. In some embodiments, the treatment is one or more differentiation induction therapies containing retinoids such as retinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapeutic agents selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group consisting of monoclonal antibody therapies such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), nonspecific immunotherapies such as BCG, interleukin-2 (IL-2), and interferon-alpha, and adjuvants, immunomodulators, e.g., thalidomide and lenalidomide (REVLIMID). Selecting an appropriate cancer therapeutic agent based on characteristics such as tumor type, cancer stage, prior exposure to cancer treatment or therapeutic agents, and other characteristics of cancer is within the capabilities of a skilled physician or oncologist.

[0334] Implementation of the VE kit Kits for carrying out the above methods, including methods related to cancer classifiers, are also disclosed herein. A kit may comprise one or more collection containers for collecting samples from an individual containing genetic material. Examples of samples include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such a kit may comprise reagents for isolating nucleic acids from the sample. The reagents may further comprise reagents for sequencing the nucleic acids, comprising buffers and detection agents. In one or more embodiments, the kit may comprise one or more sequencing panels comprising probes for targeting specific genomic regions (e.g., one or more regions found in Table 1), specific mutations, specific gene variants, or some combination thereof. In other embodiments, samples collected via the kit are provided to a sequencing laboratory where nucleic acids in the samples can be sequenced using the sequencing panels. WBC contamination detection may be applied to various configurations of the kit to minimize WBC contamination potentially originating from the kit's components. For example, experiments can be performed to compare different types of collection containers. By evaluating WBC contamination and comparing different types of collection containers, it is possible to identify the optimal type that minimizes WBC contamination.

[0335] The kit may further include instructions for using the reagents included in the kit. For example, the kit may include instructions for taking a sample and extracting nucleic acids from the test sample. Exemplary instructions may include the order in which reagents are added, the centrifugation rate used to isolate nucleic acids from the test sample, the method for amplifying nucleic acids, the method for sequencing nucleic acids, or any combination thereof. The instructions may further describe how to operate the computing device as the analytical system 200 for the purpose of performing any step of the described method.

[0336] In addition to the components described above, the kit may include a computer-readable storage medium for storing computer software for performing the various methods described throughout this disclosure. One form in which these instructions may exist is information printed on a suitable medium or substrate, such as one or more sheets of paper in the accompanying documentation within the kit's packaging. Yet another means may be computer-readable media, such as diskettes, CDs, hard drives, or network data storage in which instructions are stored in the form of computer code. Yet another means may exist may be a website address or QR code that can be used via the Internet to access information at a remote site.

[0337] Detection and mitigation of VF contamination sources In some embodiments, the method and / or classifier of the present invention is used to detect WBC contamination, for example, as described by the various methodologies in Section III.

[0338] The analytical system may utilize a WBC contamination workflow to identify the source of WBC contamination. To identify the source of contamination, the analytical system may isolate one or more variables of the sample processing workflow. The analytical system may process a first sample using a first sample processing workflow and a second set of samples using a second sample processing workflow, the second sample processing workflow having one or more variables being evaluated differently. For example, the second sample processing workflow may include a different protocol, a different clinical product, or a different sequencing device. The protocol may include steps performed in sample processing, such as centrifugation, storage temperature, and storage period. The clinical product is a manufactured product used in the sample processing workflow and may include, for example, any container used in the workflow, any chemical, any compound, any buffer, any solution, any enzyme, or any other product. The sequencing device may generally include a sequencer, but may also include other devices related to the sequencing process. The sample processing workflow may further include other laboratory devices, such as centrifuges, storage devices, and other laboratory devices for sample processing.

[0339] The analysis system applies a contamination model (e.g., one of the described contamination models) to samples from both sequencing workflows. Based on the contamination metrics for each sample set, the analysis system may determine aggregate metrics for each sequencing workflow. For example, the aggregate metrics may be a representative value, weighted average, count, or percentage of samples exceeding the contamination threshold, or a combination thereof. By comparing the aggregate contamination metrics, the analysis system can specifically identify the second sample processing workflow as contributing to WBC contamination. If a significant difference exists, the analysis system can identify the source of contamination based on which variables differ between the first and second sample processing workflows. Corrective actions may also be implemented.

[0340] Once the source of contamination is identified, the analytical system can determine the optimal sample handling workflow to mitigate WBC contamination. For example, through repeated testing, the analytical system can determine a set of protocols that minimize WBC contamination, a set of clinical products that minimize WBC contamination, a set of one or more sequencing devices that minimize WBC contamination, or a combination of these. The optimal sample handling workflow is then applied to subsequent samples.

[0341] VI. Additional Considerations The above-mentioned detailed description of embodiments refers to the accompanying drawings illustrating specific embodiments of the present disclosure. Other embodiments having different structures and operations do not depart from the scope of the present disclosure. The terms “the present invention,” etc., are used in reference to many alternative aspects or specific examples of the many alternative forms or embodiments of the applicant’s invention described herein, and neither their use nor their absence is intended to limit the scope of the applicant’s invention or the claims.

[0342] Embodiments of the present invention may also relate to apparatus for performing the operations described herein. Such apparatus may comprise a general-purpose computing device that is specifically constructed for a particular purpose and / or selectively invoked or reconfigured by a computer program stored in the computer. Such computer programs may be stored in a non-temporary, tangible, computer-readable storage medium that can be coupled to a computer system bus, or in any type of medium suitable for storing electronic instructions. Furthermore, any computing system referred to herein may comprise a single processor or may be an architecture employing multiple processor designs to enhance computing power.

[0343] Any of the steps, operations, or processes described herein as being performed by the analysis system can be performed or implemented by one or more hardware or software modules of the device alone or in combination with other computing devices. In one embodiment, the software module is implemented in a computer program product including a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.

Claims

1. A method for identifying contamination markers for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: To obtain a first set of cfDNA samples from the first cohort of subjects, and a second set of WBC samples from the same first cohort of subjects, For each sample, determine the coverage of sequence reads that overlap with each genomic locus in the initial set of genomic loci, For each genomic locus, the coverage ratio between the WBC sample and the cfDNA sample is determined, The feature set of genomic loci is determined as a contamination marker whose coverage ratio exceeds a threshold, Methods that include...

2. The method according to claim 1 or any other claim dependent thereon, wherein each cfDNA sample essentially consists of sequence reads corresponding to cfDNA fragments, and each WBC sample essentially consists of sequence reads corresponding to DNA fragments excreted from the WBC.

3. The method according to claim 1 or any other claim dependent thereon, wherein each subject of the first cohort is associated with at least one cfDNA sample and at least one WBC sample.

4. The method according to claim 1 or any other dependent claim, further comprising normalizing the coverage of sequence reads that overlap with each genomic locus based on the sequencing depth of the sample for each sample.

5. The method according to claim 1 or any other claim dependent thereon, wherein each genomic locus comprises at least one CpG site.

6. To generate the aforementioned initial set of genomic loci, the method further includes pruning one or more genomic loci with highly variable coverage across samples. The method according to claim 1 or any other claim dependent thereon.

7. Further comprising pruning one or more genomic loci with coverage that is non-normally distributed across samples, The method according to claim 1 or any other claim dependent thereon.

8. For each genomic locus, the ratio of coverage between the WBC sample and the cfDNA sample is determined. For each subject of the first cohort, the coverage ratio of each genomic locus is calculated by dividing the coverage of the genomic locus from the WBC sample by the coverage of each genomic locus from the cfDNA sample. The coverage ratio of each genomic locus is determined as a representative value of the coverage ratio of the genomic locus across the subjects of the first cohort, The method according to claim 3, including the method described in claim 3.

9. A method for training a contamination model for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: To obtain a set of cfDNA samples from non-cancer subjects, wherein each cfDNA sample contains sequence reads corresponding to cfDNA fragments, For each sample, the coverage of sequence reads that overlap with each genomic locus in the feature set of genomic loci is determined as a contamination marker, For each genomic locus in the aforementioned feature set of genomic loci, a coverage distribution is generated based on the coverage of a first subset of cfDNA samples. For each cfDNA sample in the second subset, the statistical likelihood of observing the coverage at each genomic locus is determined based on the distribution of coverage at each genomic locus, For each cfDNA sample in the second subset, contamination metrics are generated by combining the statistical likelihoods of the cfDNA samples between the genomic loci, To determine the distribution of contamination metrics for the second subset of cfDNA samples, Based on the aforementioned distribution of contamination metrics, the contamination threshold is determined, Includes, A method wherein the contamination model includes the distribution of coverage across the feature set of genomic loci, the distribution of contamination metrics, and the contamination threshold.

10. For each sample, the coverage of sequence reads that overlap with each genomic locus is further normalized based on the sequencing depth of the sample. The method according to claim 9 or any other claim dependent thereon.

11. The method according to claim 9 or any other claim dependent thereon, wherein each genomic locus comprises at least one CpG site.

12. The method according to claim 9 or any other dependent claim, wherein the set of features of a genomic locus is identified according to the method according to claim 1 or any other dependent claim.

13. The method according to claim 9 or any other claim dependent thereon, wherein the statistical likelihood for observing the coverage at each genomic locus is a z-score based on the mean and standard deviation defining the distribution of the genomic loci.

14. The method according to claim 13, wherein the impurity metric is the sum of the absolute values ​​of the z-scores.

15. The method according to claim 9 or any other claim dependent thereon, wherein the statistical likelihood for observing the coverage at each genomic locus is a p-value based on the mean and standard deviation defining the distribution of the genomic loci.

16. The method according to claim 15, wherein the contamination metric is the truncated p-value product of the p-values ​​of the cfDNA sample between the genomic loci.

17. The method according to claim 9 or any other claim dependent thereon, wherein the contamination threshold is determined to achieve a set specificity for the contamination model.

18. The method according to claim 17, wherein the specificity is one of 95%, 96%, 97%, 98%, 99%, and 99.5%.

19. A method for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: For each set of feature quantities of a genomic locus as a contamination marker, coverage is determined based on the count of sequence reads of the cfDNA fragment that overlaps with the genomic locus, For each of the feature sets of a genomic locus, the statistical likelihood of observing the coverage at that genomic locus is determined based on the distribution of coverage for that genomic locus generated from the purified cfDNA sample. By summing the statistical likelihoods across multiple contamination markers, a contamination metric for the test sample is generated. If the contamination metric for the test sample crosses the contamination threshold, it is determined whether the test sample has WBC contamination. In accordance with the determination that the test sample contains WBC contamination, one or more corrective measures shall be implemented. Methods that include...

20. The method according to claim 19 or any other claim, wherein the set of features of a genomic locus is identified according to the method according to claim 1 or any other claim dependent thereon.

21. The method according to claim 19 or any other claim thereof, wherein the distribution of the genomic loci is generated according to the method according to claim 9 or any other claim thereof.

22. The method according to claim 19 or any other claim thereof, wherein the contamination threshold is identified according to the method according to claim 9 or any other claim thereof.

23. The method according to claim 19 or any other claim dependent thereon, wherein the purified cfDNA sample essentially consists of sequence reads corresponding to cfDNA fragments.

24. The method according to claim 19 or any other claim dependent thereon, wherein the statistical likelihood for observing the coverage at each genomic locus is a z-score based on the mean and standard deviation defining the distribution of the genomic locus.

25. The method according to claim 24 or any other claim dependent thereon, wherein the impurity metric is the sum of the absolute values ​​of the z-scores.

26. The method of claim 24 or any other claim dependent thereon, wherein determining whether the contamination metric for the test sample crosses the contamination threshold includes determining whether the contamination metric exceeds the contamination threshold.

27. The method according to claim 19 or any other claim dependent thereon, wherein the statistical likelihood for observing the coverage at each genomic locus is a p-value based on the mean and standard deviation defining the distribution of the genomic locus.

28. The method according to claim 27 or any other claim dependent thereon, wherein the contamination metric is the truncated p-value product of the p-values ​​of the cfDNA sample in the genomic locus.

29. The method of claim 27 or any other claim dependent thereon, wherein determining whether the contamination metric for the test sample crosses the contamination threshold includes determining whether the contamination metric is less than the contamination threshold.

30. The aforementioned corrective measures are, To notify healthcare providers that the aforementioned test sample contains WBC contamination, Discard the aforementioned test sample. The aforementioned test sample is labeled as mixed, Notify healthcare providers to collect subsequent test samples. Notify healthcare professionals of potential sources of contamination that may contain WBCs, and Retaining the aforementioned test samples from downstream analyses, which may include cancer classification. The method according to claim 19 or any other claim dependent thereon, including any combination thereof.

31. A method for training a deconvolution model configured to deconvolve the tissue type contribution rate in a cell-free DNA (cfDNA) sample, wherein the method is The method involves obtaining a set of cfDNA samples from different tissue types, wherein each cfDNA sample contains methylated sequence reads corresponding to cfDNA fragments. For each cfDNA sample, the methylation features at each genomic locus in the initial set of genomic loci are determined based on the methylation sequence reads, For each sample, a feature vector is generated based on the methylation features across the initial set of genomic loci, To train a deconvolution model to predict the proportion of tissue types based on the feature vectors from the aforementioned sample, Methods that include...

32. Based on the information gain of the methylation features of each genomic locus based on the trained deconvolution model, the feature set of the genomic locus is identified. Based on the aforementioned set of features of the genomic loci, the deconvolution model is retrained with the updated feature vector, The method according to claim 31 or any other claim dependent thereon, further comprising:

33. The method according to claim 31 or any other claim subordinate thereto, wherein the tissue type comprises the group consisting of erythrocyte progenitor cell type, megakaryocyte cell type, monocyte cell type, granulocyte cell type, lymphocyte cell type, epithelial cell type, vascular endothelial cell type, hepatocyte type, adipocyte type, endocrine cell type, and myocyte type.

34. The method according to claim 31 or any other claim dependent thereon, wherein the tissue type comprises the group consisting of circulating bone marrow tissue type, non-circulating bone marrow tissue type, lymphocyte cell type, epithelial cell type, vascular endothelial cell type, and hepatocyte type.

35. The method according to claim 31, wherein each cfDNA sample is purified.

36. The method according to claim 31 or any other claim dependent thereon, wherein each genomic locus covers at least one CpG site.

37. The methylation features at each genomic locus are, The methylation density across the methylated sequence reads of the cfDNA sample at the aforementioned genomic locus, The count or percentage of methylated sequence reads in the cfDNA sample that are highly methylated and overlap with the genomic locus, The count or percentage of methylated sequence reads of the cfDNA sample that are in a highly demethylated state and overlap with the genomic locus, The count or percentage of methylated sequence reads having a specific methylation variant at the aforementioned genomic locus. The method according to claim 31 or any other claim dependent thereon, which is one of those claims.

38. The method according to claim 31 or any other claim dependent thereon, wherein the deconvolution model is a machine learning model.

39. The method according to claim 31 or any other claim dependent thereon, wherein the trained deconvolution model is configured to output predictions of tissue type proportions, where each tissue type proportion is the percentage of tissue type contribution to the cfDNA fragments in the cfDNA sample.

40. The method according to claim 31 or any other claim dependent thereon, wherein the information gain indicates the discriminative power of the methylation feature of each genomic locus in distinguishing between two tissue types.

41. A method for training a contamination model for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: The method involves obtaining a set of cfDNA samples from non-cancer subjects, wherein each cfDNA sample contains methylated sequence reads corresponding to cfDNA fragments. For each cfDNA sample, the methylation features at each genomic locus in the set of genomic locus features are determined based on the methylation sequence reads, For each cfDNA sample, a feature vector is generated based on the methylation features across the set of features of the genomic locus, In order to predict the proportion of tissue types for the aforementioned cfDNA samples, a trained deconvolution model is applied to the feature vectors of each cfDNA sample, To construct a distribution of the proportion of the tissue types for each of the aforementioned cfDNA samples, A method comprising the aforementioned distribution, wherein the contamination model includes the aforementioned distribution.

42. The method according to claim 41 or any other claim according to claim 41, wherein the set of features of a genomic locus is identified and / or the deconvolution model is trained, according to the method according to claim 31 or any other claim dependent thereon.

43. The method according to claim 41 or any other claim dependent thereon, wherein each cfDNA sample is purified.

44. The method according to claim 41 or any other claim dependent thereon, wherein each genomic locus covers at least one CpG site.

45. The methylation features at each genomic locus are The methylation density across the methylated sequence reads of the cfDNA sample at the aforementioned genomic locus, The count or percentage of methylated sequence reads in the cfDNA sample that are highly methylated and overlap with the genomic locus, The count or percentage of methylated sequence reads of the cfDNA sample that are in a highly demethylated state and overlap with the genomic locus, The count or percentage of methylated sequence reads having a specific methylation variant at the aforementioned genomic locus. The method according to claim 41 or any other claim dependent thereon, which is one of those claims.

46. The method according to claim 41 or any other claim dependent thereon, wherein the distribution is a multivariate distribution in a multivariate vector space where the number of variables is equal to the number of organizational types.

47. A method for detecting the presence of leukocytes (WBCs) in a test sample containing methylated sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: Based on the methylation sequence reads, the methylation features at each genomic locus in the feature set of genomic loci are determined, To generate a feature vector based on the methylation features across the feature set of genomic loci, In order to predict the proportion of tissue types in the aforementioned test sample, a trained deconvolution model is applied to the feature vector of the aforementioned test sample, The method involves generating a contamination metric for the test sample based on the distance of the proportion of tissue types in the test sample to the distribution of the proportion of tissue types generated from cfDNA samples from non-cancer subjects, wherein the contamination metric indicates the likelihood that the cfDNA sample has WBC contamination. If the contamination metric for the test sample crosses the contamination threshold, it is determined whether the test sample has WBC contamination. In accordance with the determination that the test sample contains WBC contamination, one or more corrective measures shall be implemented. Methods that include...

48. The method according to claim 47 or any other claim according to claim 47, wherein the set of features of a genomic locus is identified and / or the deconvolution model is trained according to the method according to claim 31 or any other claim dependent thereon.

49. The method according to claim 47 or any other claim thereof, wherein the deconvolution model is trained according to the method according to claim 41 or any other claim thereof.

50. The method according to claim 47 or any other claim dependent thereon, wherein each genomic locus covers at least one CpG site.

51. The methylation features at each genomic locus are The methylation density across the methylated sequence reads of the cfDNA sample at the aforementioned genomic locus, The count or percentage of methylated sequence reads in the cfDNA sample that are highly methylated and overlap with the genomic locus, The count or percentage of methylated sequence reads of the cfDNA sample that are in a highly demethylated state and overlap with the genomic locus, The count or percentage of methylated sequence reads having a specific methylation variant at the aforementioned genomic locus. The method according to claim 47 or any other claim dependent thereon, which is one of those claims.

52. The method according to claim 47 or any other claim dependent thereon, wherein the distance is a Mahalanobis distance based on the distribution of tissue types.

53. The method according to claim 47 or any other claim dependent thereon, wherein the contamination metric is a p-value.

54. The method according to claim 53, wherein determining whether the contamination metric for the test sample crosses the contamination threshold includes determining whether the contamination metric is less than the contamination threshold.

55. The method according to claim 47 or any other claim dependent thereon, wherein the contamination threshold is determined to achieve a set specificity for the contamination model.

56. The aforementioned corrective measures are, To notify healthcare providers that the aforementioned test sample contains WBC contamination, Discard the aforementioned test sample. The aforementioned test sample is labeled as mixed, Notify healthcare providers to collect subsequent test samples. Notify healthcare professionals of potential sources of contamination that may contain WBCs, and Retaining the aforementioned test samples from downstream analyses, which may include cancer classification. The method according to claim 47 or any other claim dependent thereon, including any combination thereof.

57. A method for training a contamination model for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: The objective is to obtain a first set of cfDNA samples and a second set of WBC samples, wherein each sample contains sequence reads corresponding to DNA fragments. For each sample in the first set and the second set, the average coverage of sequence reads that overlap with the initial set of genomic loci is determined, For each sample in the first set and the second set, the normalized coverage for each genomic locus in the initial set of genomic loci is determined based on sequence reads that overlap with the genomic loci normalized by the average coverage of the sample, For each genomic locus, a first coverage distribution for cfDNA samples and a second coverage distribution for WBC samples are generated. Based on the aforementioned coverage distribution, the goal is to identify highly distinctive genomic loci between cf DNA samples and WBC samples, The discriminative score for each genomic locus is determined using a two-sample t-test, Based on the aforementioned discriminative score, a set of feature quantities for genomic loci is determined as a contamination marker, Methods that include...

58. The method according to claim 57 or any other claim dependent thereon, wherein each cfDNA sample essentially consists of sequence reads corresponding to cfDNA fragments, and each WBC sample essentially consists of sequence reads corresponding to DNA fragments discharged from the WBC.

59. The method according to claim 57 or any other claim dependent thereon, wherein each genomic locus comprises at least one CpG site.

60. To generate the aforementioned initial set of genomic loci, pruning one or more genomic loci with highly variable coverage across samples. The method according to claim 57 or any other claim dependent thereon, further comprising:

61. Pruning one or more genomic loci with coverage that is non-normally distributed across samples. The method according to claim 57 or any other claim dependent thereon, further comprising:

62. Based on the aforementioned coverage distribution, it is possible to identify highly discriminative genomic loci between cfDNA samples and WBC samples. For each genomic locus, evaluate the distance between the first distribution of coverage for cfDNA samples and the second distribution of coverage for WBC samples. The method according to claim 57 or any other claim dependent thereon, including the method described therein.

63. The method according to claim 62, wherein the distance for each genomic locus is based on the area under the curve (AUC) for the binary classification between coverage-based cfDNA and WBC at the genomic locus.

64. The method according to claim 63, wherein the distance for each genomic locus is further based on the Hodges-Lehmann estimator.

65. The method according to claim 57 or any other claim dependent thereon, wherein the two-sample t-test evaluates the discriminability of paired cfDNA samples and WBC samples for each subject.

66. The method according to claim 57 or any other claim dependent thereon, wherein the distinctive score is based on the Bonferroni correction.

67. The method according to claim 57 or any other claim dependent thereon, further comprising pruning any genomic loci having a non-normal distribution of coverage for cfDNA samples or a non-normal distribution of coverage for WBC samples.

68. The method according to claim 57 or any other claim dependent thereon, further comprising pruning any genomic loci having a difference below a threshold between the mean of the distribution of coverage for cfDNA samples and the mean of the distribution of coverage for WBC samples.

69. Based on the aforementioned discriminative score, the set of features of the genomic locus is determined as a contamination marker. Identifying genomic loci with discriminative scores exceeding a threshold, The method according to claim 57 or any other claim dependent thereon, including the method described therein.

70. Based on the aforementioned discriminative score, the set of features of the genomic locus is determined as a contamination marker. Ranking the initial set of genomic loci based on the aforementioned discriminative score, In order to form the aforementioned set of features for genomic loci, the top number of genomic loci are selected from the aforementioned ranking, The method according to claim 57 or any other claim dependent thereon, including the method described therein.

71. A method for detecting the presence of leukocytes (WBCs) in a test sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: Determining the average coverage of sequence reads that overlap with the feature set of genomic loci, Based on sequence reads that overlap with the genomic loci normalized by the average coverage of the sample, the normalized coverage for each genomic locus in the feature set of the genomic loci is determined. A contamination model for determining contamination metrics as the contribution rate of DNA excreted by the WBC is applied to the test sample that maximizes the likelihood of observing the normalized coverage across the feature set, based on the coverage distribution for the cfDNA sample and the WBC sample for the feature set of the genomic locus, If the contamination metric for the test sample crosses the contamination threshold, it is determined whether the test sample has WBC contamination. In accordance with the determination that the test sample contains WBC contamination, one or more corrective measures shall be implemented. Methods that include...

72. The method according to claim 71 or any other claim according to claim 57 or any other claim dependent thereon, wherein the set of features of a genomic locus is identified, the contamination model is trained, and / or the distribution of coverage for cfDNA and WBC samples is generated.

73. The method according to claim 71 or any other claim dependent thereon, wherein the contamination model includes, for each genomic locus, a first distribution of coverage for cfDNA samples and a second distribution of coverage for WBC samples.

74. The method of claim 71 or any other claim dependent thereon, wherein determining whether the contamination metric for the test sample crosses the contamination threshold includes determining whether the contamination metric is less than the contamination threshold.

75. The method according to claim 71 or any other claim dependent thereon, wherein the contamination threshold is determined to achieve a set specificity for the contamination model.

76. The aforementioned corrective measures are, To notify healthcare providers that the aforementioned test sample contains WBC contamination, Discard the aforementioned test sample. The aforementioned test sample is labeled as mixed, Notify healthcare providers to collect subsequent test samples. Notify healthcare professionals of potential sources of contamination that may contain WBCs, and Retaining the aforementioned test samples from downstream analyses, which may include cancer classification. The method according to claim 71 or any other claim dependent thereon, including any combination thereof.

77. A method for determining the source of leukocyte (WBC) contamination in one or more samples containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is: To obtain a first set of samples from the first sample processing workflow, The method involves obtaining a second set of samples from a second sample processing workflow, wherein the second sample processing workflow includes a first protocol different from the first sample processing workflow, a first clinical product different from the first sample processing workflow, a first sequencing device, or a combination thereof. Applying a contamination model to the first set of samples and the second set of samples in order to determine contamination metrics for each sample, Based on the combination of contamination metrics for the first set of samples, a first aggregate metric for the first sample processing workflow is determined, Based on the combination of contamination metrics for the second set of samples, a second aggregate metric for the second sample processing workflow is determined, Based on the determination that the second aggregate metric is greater than the first aggregate metric, it is determined that the source of the WBC contamination originates from the first protocol, the first clinical product, the first sequencing device, or a combination thereof. One or more corrective measures are implemented from the second sample processing workflow described above to mitigate the WBC contamination, Methods that include...

78. The aforementioned contamination model, Claim 9 or any other claim dependent thereon, Claim 41 or any other claim dependent thereon, Claim 57 or any other claim dependent thereon The method of claim 77, which is trained according to the method thereof.

79. The aforementioned corrective measures are, Changing the first protocol in the second sample processing workflow to another protocol, Replacing the first clinical product in the second sample processing workflow with another clinical product, Replacing the first sequencing device in the second sample processing workflow with another sequencing device, Obtaining a third set of samples instead of the second set of sequences determined from the first sample processing workflow, and Discard the second set of samples obtained from the second sample processing workflow. The method of claim 77 or any other claim dependent thereon, including any combination thereof.

80. The method according to claim 77 or any other claim dependent thereon, wherein the first aggregated contamination metric is a representative value of the contamination metrics of a first set of samples, and the second aggregated contamination metric is a representative value of the contamination metrics of a second set of samples.

81. The method according to claim 77 or any other claim dependent thereon, wherein the first aggregated contamination metric is a first proportion or first count of samples in a first set of samples having a corresponding contamination metric exceeding a contamination threshold, and the second aggregated contamination metric is a second proportion or second count of samples in a second set of samples having a corresponding contamination metric exceeding a contamination threshold.

82. A method for training a cancer classifier using a sample containing sequence reads corresponding to cell-free DNA (cfDNA) fragments, wherein the method is To obtain a first set of samples from a first cohort of healthy subjects and a second set of samples from a second cohort of subjects diagnosed with cancer, Applying a contamination model to the first set of samples and the second set of samples to determine contamination metrics for each sample, which indicate the amount of leukocyte (WBC) contamination in the aforementioned sample, Identify one or more samples with contamination conditions that have corresponding contamination metrics exceeding the contamination threshold, Filtering the contaminated sample, and the filtering may include discarding the contaminated sample. Based on the sequence reads of the sample, the feature vectors for each remaining sample are determined, The cancer classifier is trained using the feature vectors of the remaining sample, wherein the trained cancer classifier is configured to predict the likelihood of cancer presence based on input feature vectors derived from sequence reads in the test cfDNA sample. Methods that include...

83. The aforementioned contamination model, Claim 9 or any other claim dependent thereon, Claim 41 or any other claim dependent thereon, Claim 57 or any other claim dependent thereon The method of claim 82, wherein the person is trained according to the method described in [the method].

84. The method according to claim 82 or any other claim dependent thereon, wherein the cancer classifier is a machine learning model.

85. The method according to claim 82 or any other claim dependent thereon, wherein the cancer classifier is trained to predict a binary prediction between the presence or absence of cancer.

86. The method according to claim 82 or any other claim dependent thereon, wherein the second set of samples comprises one or more subsets of samples, each subset of samples obtained from a subject diagnosed with one of a plurality of cancer types.

87. The method according to claim 86, wherein the cancer classifier is trained to predict multi-class predictions as the likelihood of the presence of multiple types of cancer.

88. The method according to claim 82 or any other claim dependent thereon, wherein the sequence reads for each sample are methylated sequence reads, and the feature vector is based on methylated sequence reads having a highly informative methylation pattern.

89. A non-temporary computer-readable storage medium that stores instructions causing one or more computer processors to perform the method according to any one of claims 1 to 88 when executed by one or more computer processors.

90. It is a system, One or more computer processors, A non-temporary computer-readable storage medium according to claim 89, A system that includes this.

91. It is a treatment kit, A collection container for collecting DNA samples from the subject, One or more reagents for optionally isolating DNA fragments from a DNA sample, One or more contaminating probes targeting one or more genomic loci selected at will from Table 1, A non-temporary computer-readable storage medium according to claim 89, A treatment kit that includes this.