Methods and systems for detecting colorectal cancer by nucleic acid methylation analysis

By combining a methylation signature panel and a machine learning classifier, the shortcomings of existing colorectal cancer screening tools in terms of sensitivity and specificity are addressed, enabling efficient screening for early detection and monitoring of colorectal cancer, and supporting the detection of treatment response and minimal residual disease.

CN115667554BActive Publication Date: 2026-02-27FREENOM HLDG INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180039398.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-31
Filing Date
2021-03-29
Publication Date
2026-02-27
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Existing cancer screening tools have trade-offs in false positive and false negative results in colorectal cancer detection, and lack sufficient sensitivity and specificity, making it difficult to detect early signs of colorectal cancer with low tumor burden or recurrence, especially in the initial screening of high-risk groups.

Method used

By employing a methylation signature panel and a machine learning classifier, this study analyzes the genomic methylation profiles of individual biological samples to identify methylation signatures specific to colonic proliferative disorders. The results are then combined with a machine learning model for classification, enabling early detection and monitoring of colorectal cancer.

Benefits of technology

It improves the sensitivity and specificity of colorectal cancer detection, enabling early detection of colorectal cancer in high-risk groups, reducing false positive rates, providing more accurate screening tools, and supporting monitoring of treatment response and detection of minimal residual disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115667554B_ABST
    Figure CN115667554B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and systems for screening or detecting colorectal cancer or subsequent colorectal disease progression, which can be applied to cell-free nucleic acids, such as cell-free DNA. The methods can train a machine learning model using detection of methylation signals within single sequencing reads in identified genomic regions as input features, and generate a classifier suitable for stratifying a population of individuals. The methods can include extracting DNA from a cell-free sample obtained from a subject, transforming the DNA for methylation sequencing, generating sequencing reads, and detecting colon proliferative cell disorder related signals in the sequencing information, and training a machine learning model to provide a discriminator capable of distinguishing between groups such as healthy, cancer, etc. in a population of subjects, or distinguishing between disease subtypes or stages. The methods can be used, for example, to predict, prognosticate, and / or monitor response to treatment, tumor burden, recurrence, or progression of colorectal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 002,878, filed March 31, 2020, the contents of which are hereby incorporated by reference in their entirety. BACKGROUND

[0003] The present disclosure relates generally to cancer detection and disease monitoring. More specifically, the field relates to cancer-related DNA methylation detection and disease monitoring of early colorectal cancer (CRC). Over the past few decades, cancer screening and monitoring can help improve outcomes, as early detection leads to better outcomes, so that cancer can be eliminated before it spreads. For example, in the case of CRC, the use of colonoscopy can play a role in improving early diagnosis. Unfortunately, there can be some challenges due to patient compliance with screening not being as regular as recommended.

[0004] A major problem with any screening tool can be the trade-off between false positive and false negative results (or specificity and sensitivity), leading to unnecessary investigations in the former case and inefficiencies in the latter. An ideal test can be one with a high positive predictive value (PPV), minimizing unnecessary investigations, but detecting the vast majority of cancers. Another key factor can be the so-called “detection sensitivity” and also the lower limit of tumor size for detection, which is distinguished from test sensitivity. Unfortunately, waiting for the tumor to grow large enough to release circulating tumor markers at the necessary level for detection can be at odds with the requirement for early detection, which is to treat the tumor at a stage where treatment is most effective. Thus, there is a need for effective blood-based screening for early CRC based on circulating analytes.

[0005] Detection of circulating tumor DNA is increasingly being recognized as a viable “liquid biopsy” allowing for non-invasive detection and information investigation of tumors. In some cases, these techniques have been applied to colon, breast, and prostate cancer through identification of tumor-specific mutations. Sensitivity of these techniques can be limited due to the high background of normal (e.g., non-tumor-derived) DNA present in circulation.

[0006] Detection of tumor-specific methylation in blood can provide a distinct advantage over mutation detection. A number of single or multiple methylation biomarkers can be assessed in cancers including lung, colon, and breast cancer. These can have low sensitivity as they can not be sufficiently prevalent in tumors.

[0007] There remains a need for more sensitive and specific screening tools to detect early or low tumor burden colorectal cancer tumor signals in recurrence and for primary screening in high-risk populations. SUMMARY

[0008] The present disclosure provides methods and systems related to gene methylation profiling associated with colorectal cancer detection and disease progression.

[0009] In one aspect, the present disclosure provides a methylation signature panel specific for a colon cell proliferative disorder comprising: one or more methylation genomic regions selected from Table 11, wherein the one or more regions are more highly methylated in a biological sample from an individual having a colon cell proliferative disorder or a subtype of a colon cell proliferative disorder and less methylated in normal tissue and normal blood cells from an individual not having a colon cell proliferative disorder.

[0010] In some embodiments, the biological sample is a nucleic acid, DNA, ribonucleic acid (RNA), or cell-free nucleic acid (e.g., cfDNA or cfRNA).

[0011] In some embodiments, the genomic regions are classified as non-coding, coding, or non-transcribed or regulatory regions.

[0012] In some embodiments, the signature panel comprises increased methylation in two or more genomic regions selected from Table 11.

[0013] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of: cell-free DNA, cell-free RNA, bodily fluid, stool, colonic discharge, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0014] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal cancer.

[0015] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, or stage 4 colorectal cancer.

[0016] In some embodiments, the signature panel comprises two or more methylation genomic regions in Table 1-11, three or more methylation genomic regions in Table 1-11, four or more methylation genomic regions in Table 1-11, five or more methylation genomic regions in Table 1-11, six or more methylation genomic regions in Table 1-11, seven or more methylation genomic regions in Table 1-11, eight or more methylation genomic regions in Table 1-11, nine or more methylation genomic regions in Table 1-11, ten or more methylation genomic regions in Table 1-11, eleven or more methylation genomic regions in Table 1-11, twelve or more methylation genomic regions in Table 1-11, or thirteen or more methylation genomic regions in Table 1-11.

[0017] In some embodiments, the signature panel comprises genomic regions that are methylated in colorectal cancer, including methylation regions in one or more genomic regions selected from the group consisting of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, and ZNF543.

[0018] In some embodiments, the regions that are methylated in colorectal cancer comprise methylation regions in the ITGA4 and EMBP1 genomic regions.

[0019] In some embodiments, the regions that are methylated in colorectal cancer comprise methylation regions in one or more genomic regions selected from the group consisting of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.

[0020] In some embodiments, the signature panel comprises methylation genomic regions selected from Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, and Table 11.

[0021] In another aspect, the present disclosure provides a methylation signature panel specific to a colon cell proliferative disorder, comprising: two or more methylation genomic regions in Table 1-11, wherein the two or more regions are more highly methylated in a biological sample from an individual having a colon cell proliferative disorder or a subtype of a colon cell proliferative disorder, and less methylated in normal tissue and normal blood cells of an individual not having a colon cell proliferative disorder.

[0022] In some embodiments, the biological sample is a nucleic acid, DNA, ribonucleic acid (RNA), or cell-free nucleic acid (cfDNA or cfRNA).

[0023] In some embodiments, the genomic regions are divided into non-coding regions, coding regions, or non-transcribed regions or regulatory regions.

[0024] In some embodiments, the signature panel comprises increased methylation in 6 or more or 12 or more genomic regions in Tables 1-11.

[0025] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of: cell-free DNA, cell-free RNA, bodily fluid, stool, colonic discharge, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0026] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial cancer, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal cancer.

[0027] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, or stage 4 colorectal cancer.

[0028] In some embodiments, the signature panel comprises three or more methylation genomic regions in Tables 1-11, four or more methylation genomic regions in Tables 1-11, five or more methylation genomic regions in Tables 1-11, six or more methylation genomic regions in Tables 1-11, seven or more methylation genomic regions in Tables 1-11, eight or more methylation genomic regions in Tables 1-11, nine or more methylation genomic regions in Tables 1-11, ten or more methylation genomic regions in Tables 1-11, eleven or more methylation genomic regions in Tables 1-11, twelve or more methylation genomic regions in Tables 1-11, or thirteen or more methylation genomic regions in Tables 1-11.

[0029] In some embodiments, the signature panel comprises genomic regions that are methylated in colorectal cancer, including methylation regions in one or more genomic regions selected from the group consisting of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, and ZNF543.

[0030] In some embodiments, the regions that are methylated in colorectal cancer comprise methylation regions in the ITGA4 and EMBP1 genomic regions.

[0031] In some embodiments, the regions methylated in colorectal cancer include methylation regions in one or more genomic regions selected from the group consisting of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.

[0032] In some embodiments, the signature panel comprises methylation regions selected from Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11.

[0033] In another aspect, the present disclosure provides a classifier (e.g., a machine learning classifier) capable of distinguishing between a population of healthy individuals and a population of individuals having a colon cell proliferative disorder, comprising: a) a set of measurements representing differentially methylated genomic regions, wherein the measurements are obtained from methylation sequencing data from healthy subjects and subjects having a colon cell proliferative disorder; b) wherein the measurements are used to generate a set of features corresponding to properties of the differentially methylated genomic regions, and inputting the features into a machine learning or statistical model; and c) wherein the model provides a feature vector that serves as a classifier capable of distinguishing between a population of healthy individuals and a population of individuals having a colon cell proliferative disorder.

[0034] In some embodiments, the set of measurements describe features of methylation regions selected from the group consisting of: percent methylation per base of CpG, CHG, CHH, counts or ratios of fragments with different counts or ratios of methylated CpGs observed in a region, conversion efficiency (100 - average percent methylation of CHH), hypomethylated segments, methylation levels (overall average methylation of CpG, CHH, CHG, fragment length, fragment midpoint, and methylation levels in one or more genomic regions such as chrM, LINE1, or ALU), number of methylated CpGs per fragment, fraction of CpG methylation per fragment, fraction of CpG methylation per region, fraction of CpG methylation in panel, dinucleotide coverage (normalized dinucleotide coverage), coverage evenness (unique CpG sites under 1x and 10x average genome coverage (for S4 runs)), overall average CpG coverage (depth), and average coverage at CpG islands, CGI shores, and CGI shores.

[0035] In some embodiments, the machine learning model comprises a classifier loaded into a memory of a computer system, a machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having a colon cell proliferative disorder, and a second subset of the training biological samples identified as not having a colon cell proliferative disorder.

[0036] In some embodiments, the classifier is provided in a system for detecting a colon cell proliferative disorder, the system comprising: a) a computer readable medium comprising a classifier operable to classify a subject as having a colon cell proliferative disorder or not having a colon cell proliferative disorder according to a methylation signature panel; and b) one or more processors for executing instructions stored on the computer readable medium.

[0037] In some embodiments, the system comprises a classification loop configured as a machine learning classifier selected from the group consisting of: a deep learning classifier, a neural network classifier, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and a principal component analysis classifier.

[0038] In some embodiments, the computer readable medium is a non-transitory computer readable medium comprising machine executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0039] In some embodiments, the system comprises one or more computer processors and a computer memory coupled with the same. The computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any of the methods described herein.

[0040] In another aspect, the present disclosure provides a method for determining a methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample from an individual, comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules of the cfDNA sample to produce a plurality of converted nucleic acids; b) contacting the plurality of converted nucleic acids with nucleic acid probes complementary to a pre-identified methylation signature panel selected from at least two differentially methylated regions of Tables 1-11 to enrich for sequences corresponding to the signature panel; c) determining nucleic acid sequences of the plurality of converted nucleic acid molecules; and d) aligning the nucleic acid sequences of the plurality of converted nucleic acid molecules to reference nucleic acid sequences, thereby determining the methylation profile of the individual.

[0041] In some embodiments, the nucleic acid sequencing library is prepared prior to amplification. In some embodiments, the method further comprises amplifying the plurality of transformed nucleic acids. In some embodiments, the amplification comprises polymerase chain reaction (PCR). In some embodiments, the method further comprises determining the nucleic acid sequence of the transformed nucleic acid molecules at a depth greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x. In some embodiments, the reference nucleic acid sequence is at least a portion of a human reference genome. In some embodiments, the human reference genome is hg18.

[0042] In some embodiments, the methylation profile is associated with a colon cell proliferative disorder and provides a classification of the subject with respect to having a colon cell proliferative disorder.

[0043] In some embodiments, the nucleic acid adaptors comprising unique molecular identifiers are attached to the untransformed nucleic acids in the cfDNA sample prior to a).

[0044] In some embodiments, the nucleic acid molecules are subjected to cytosine to uracil conversion conditions using chemical methods, enzymatic methods, or a combination thereof.

[0045] In some embodiments, the cfDNA in the biological sample is treated with a reagent selected from the group consisting of bisulfite, hydrogen sulfite, disulfide, and combinations thereof.

[0046] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of cell-free DNA, cell-free RNA, bodily fluid, stool, colon discharge, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0047] In some embodiments, the method comprises comparing the methylation signature panel measured from the subject to a database of methylation signature panels measured from normal subjects, wherein the database is stored in a computer system; determining that the subject has an increased risk of having a colon cell proliferative disorder by measuring at least a 1%, at least a 2%, at least a 3%, at least a 4%, at least a 5%, at least a 6%, at least a 7%, at least an 8%, at least a 9%, at least a 10%, at least an 11%, at least a 12%, at least a 13%, at least a 14%, at least a 15%, at least a 16%, at least a 17%, at least a 18%, at least a 19%, or at least a 20% change in the methylation status of the methylation signature panel compared to the methylation status from normal subjects.

[0048] In some embodiments, the prequalification methylation signature panel comprises three or more methylation genomic regions in Tables 1-11, four or more methylation genomic regions in Tables 1-11, five or more methylation genomic regions in Tables 1-11, six or more methylation genomic regions in Tables 1-11, seven or more methylation genomic regions in Tables 1-11, eight or more methylation genomic regions in Tables 1-11, nine or more methylation genomic regions in Tables 1-11, ten or more methylation genomic regions in Tables 1-11, eleven or more methylation genomic regions in Tables 1-11, twelve or more methylation genomic regions in Tables 1-11, or thirteen or more methylation genomic regions in Tables 1-11. In some embodiments, the prequalification methylation signature panel comprises one or more methylation genomic regions in Table 11, two or more methylation genomic regions in Table 11, or three methylation genomic regions in Table 11. In some embodiments, the methylation profile is indicative of the presence or absence of a colon cell proliferative disorder in the individual.

[0049] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of adenoma (adenomatous polyps), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal carcinoma, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal carcinoma.

[0050] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of stage 1 colorectal carcinoma, stage 2 colorectal carcinoma, stage 3 colorectal carcinoma, and stage 4 colorectal carcinoma.

[0051] In another aspect, the present disclosure provides a method for detecting the presence or absence of a colon cell proliferative disorder in a subject, comprising: a) providing conditions capable of converting unmethylated cytosine to uracil in nucleic acid molecules of a biological sample obtained or derived from the subject to produce a plurality of converted nucleic acids; b) contacting the plurality of converted nucleic acids with nucleic acid probes complementary to a prequalification methylation signature panel of at least two differentially methylated regions selected from Tables 1-11 to enrich for sequences corresponding to the signature panel; c) determining nucleic acid sequences of the plurality of converted nucleic acid molecules; d) aligning the nucleic acid sequences of the plurality of converted nucleic acid molecules to reference nucleic acid sequences, thereby determining a methylation profile of the individual; and e) applying a trained machine learning model to the methylation profile, wherein the trained machine learning model is trained to be capable of distinguishing between a healthy individual and an individual having a colon cell proliferative disorder to provide an output value associated with the presence of a colon cell proliferative disorder, thereby detecting the presence or absence of a colon cell proliferative disorder in the subject.

[0052] In some embodiments, the nucleic acid sequencing library is prepared prior to amplification. In some embodiments, the method further comprises amplifying the plurality of transformed nucleic acids. In some embodiments, the amplification comprises polymerase chain reaction (PCR). In some embodiments, the method further comprises determining the nucleic acid sequence of the transformed nucleic acid molecules at a depth greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x. In some embodiments, the reference nucleic acid sequence is at least a portion of a human reference genome. In some embodiments, the human reference genome is hg18.

[0053] In some embodiments, the biological sample obtained from the subject is selected from the group consisting of: cell-free DNA, cell-free RNA, bodily fluid, stool, colonic discharge, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0054] In some embodiments, the method comprises comparing the methylation signature panel measured from the subject to a database of methylation signature panels measured from normal subjects, wherein the database is stored in a computer system; determining that the subject has an increased risk of developing a colon cell proliferative disorder by measuring at least a 1%, at least a 2%, at least a 3%, at least a 4%, at least a 5%, at least a 6%, at least a 7%, at least a 8%, at least a 9%, at least a 10%, at least a 11%, at least a 12%, at least a 13%, at least a 14%, at least a 15%, at least a 16%, at least a 17%, at least a 18%, at least a 19%, or at least a 20% change in the methylation status of the methylation signature panel compared to the methylation status from normal subjects.

[0055] In some embodiments, the prequalification methylation signature panel comprises three or more methylation genomic regions in Tables 1-11, four or more methylation genomic regions in Tables 1-11, five or more methylation genomic regions in Tables 1-11, six or more methylation genomic regions in Tables 1-11, seven or more methylation genomic regions in Tables 1-11, eight or more methylation genomic regions in Tables 1-11, nine or more methylation genomic regions in Tables 1-11, ten or more methylation genomic regions in Tables 1-11, eleven or more methylation genomic regions in Tables 1-11, twelve or more methylation genomic regions in Tables 1-11, or thirteen or more methylation genomic regions in Tables 1-11. In some embodiments, the prequalification methylation signature panel comprises one or more methylation genomic regions in Table 11, two or more methylation genomic regions in Table 11, or three methylation genomic regions in Table 11. In some embodiments, the methylation profile is indicative of the presence or absence of a colon cell proliferative disorder in the individual. In some embodiments, the method further comprises administering a treatment for a colon cell proliferative disorder to the individual based on detecting the presence of a colon cell proliferative disorder in the individual.

[0056] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal carcinoma, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal carcinoma.

[0057] In some embodiments, the trained machine learning classifier is selected from the group consisting of: a deep learning classifier, a neural network classifier, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and a principal component analysis classifier.

[0058] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: stage 1 colorectal carcinoma, stage 2 colorectal carcinoma, stage 3 colorectal carcinoma, and stage 4 colorectal carcinoma.

[0059] In another aspect, the disclosure provides a method for monitoring minimal residual disease in a subject previously treated for a disease, comprising: determining the methylation profile described herein as a baseline methylation state, and repeating the analysis to determine the methylation profile at one or more predetermined time points, wherein a change compared to the baseline indicates a change in the subject's minimal residual disease status at baseline.

[0060] In some embodiments, the minimal residual disease is selected from the group consisting of response to treatment, tumor burden, post-surgical residual tumor, relapse, secondary screening, primary screening, and cancer progression.

[0061] In another aspect, a method for determining a response to treatment is provided.

[0062] In another aspect, a method for monitoring tumor burden is provided.

[0063] In another aspect, a method for detecting post-surgical residual tumor is provided.

[0064] In another aspect, a method for detecting relapse is provided.

[0065] In another aspect, a method for use as secondary screening is provided.

[0066] In another aspect, a method for use as primary screening is provided.

[0067] In another aspect, a method for monitoring cancer progression is provided.

[0068] In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 80% sensitivity. In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 90% sensitivity. In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 95% sensitivity. In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 70% positive predictive value (PPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 80% positive predictive value (PPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 90% positive predictive value (PPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 95% positive predictive value (PPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 99% positive predictive value (PPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 80% negative predictive value (NPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 90% negative predictive value (NPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 95% negative predictive value (NPV). In some embodiments, the data set indicates the presence or susceptibility of colorectal cancer with at least about 99% negative predictive value (NPV). In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in the subject with an area under the curve (AUC) of at least about 0.90. In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in the subject with an area under the curve (AUC) of at least about 0.95. In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in the subject with an area under the curve (AUC) of at least about 0.99.

[0069] In some embodiments, the method further comprises displaying a report on a graphical user interface of an electronic device of a user. In some embodiments, the user is the subject, individual, or patient.

[0070] In some embodiments, the method further comprises determining a likelihood of the presence or susceptibility of colorectal cancer in the subject, individual, or patient. For example, the likelihood can be a probability value between 0% and 100%.

[0071] In some embodiments, the trained algorithm (e.g., machine learning model or classifier) comprises a supervised machine learning algorithm. In some embodiments, the supervised machine learning algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.

[0072] In some embodiments, the method further comprises providing the subject with a treatment intervention based at least in part on the methylation profile or analysis, such as a treatment intervention (e.g., chemotherapy, radiation therapy, immunotherapy, or surgery) to treat the colorectal cancer patient.

[0073] In some embodiments, the method further comprises monitoring the presence or predisposition to colorectal cancer, wherein the monitoring comprises assessing the presence or predisposition to colorectal cancer in the subject at a plurality of time points, wherein the assessment is based at least on the presence or predisposition to colorectal cancer determined at each of the plurality of time points.

[0074] In some embodiments, a difference in the assessment of the presence or predisposition to colorectal cancer in the subject at a plurality of time points is indicative of one or more clinical indications selected from the group consisting of: (i) diagnosis of the presence or predisposition to colorectal cancer in the subject, (ii) prognosis of the presence or predisposition to colorectal cancer in the subject, and (iii) efficacy or ineffectiveness of a course of treatment to treat the presence or predisposition to colorectal cancer in the subject.

[0075] In some embodiments, the method further comprises stratifying the colorectal cancer of the subject from a plurality of different colorectal cancer subtypes or stages to determine a colorectal cancer subtype of the subject by using a trained algorithm.

[0076] Another aspect of the disclosure provides a non-transitory computer readable medium comprising machine executable code that, when executed by one or more computer processors, implements any of the methods above or elsewhere herein.

[0077] Another aspect of the disclosure provides a system comprising one or more computer processors and computer memory coupled with the same. The computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0078] Additional aspects and advantages of the disclosure will be apparent from the following detailed description of the disclosure, which shows and describes only the preferred aspects of the disclosure. As will be realized, the disclosure is capable of other and different aspects and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0079] INCORPORATION BY REFERENCE

[0080] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the extent such disclosure is not inconsistent with this specification. To the extent that any publication, patent, or patent application is inconsistent with this specification, the disclosure of the present specification will control. BRIEF DESCRIPTION OF DRAWINGS

[0081] Embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings. The novel features of the application are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present application will be obtained by reference to the following detailed description that sets forth illustrative

[0082] FIG. 1 A schematic is provided for programming or otherwise configuring a computer system to implement the methods provided herein, or machine learning models and classifiers.

[0083] FIG. 2 Area under the curve (AUC) curves for 4-fold cross-validation of models trained on the regions in Table 1 are provided.

[0084] FIG. 3A-FIG. 3F A series of area under the curve (AUC) curves are provided for samples of different stages of CRC trained on the classification model. FIG. 3A-FIG. 3F ROC results are shown demonstrating the ability of these differentially methylated regions (DMRs) to detect CRC and distinguish early stage cancer, including patients with stage 1 ( FIG. 3A ), stage 2 ( FIG. 3B ), stage 3 ( FIG. 3C ), stage 4 ( FIG. 3D ), missing stage ( FIG. 3E ), and all samples ( FIG. 3F ). DETAILED DESCRIPTION

[0085] While various embodiments of the application have been shown and described herein, it will be apparent to those skilled in the art that many changes, modifications, and alternatives can be made to such embodiments without departing from the application. It should be understood that various alternatives to the embodiments of the application described herein can be employed.

[0086] The present disclosure relates generally to cancer detection and disease monitoring. More specifically, the field relates to cancer-related DNA methylation detection and early colorectal cancer disease monitoring. Over the past few decades, cancer screening and monitoring can help improve outcomes, as early detection leads to better outcomes, so that cancer can be eliminated before it spreads. In the case of colorectal cancer, for example, the use of colonoscopy can play a role in improving early diagnosis. Unfortunately, there can be some challenges due to patient compliance with screening not being as regular as recommended.

[0087] A major problem with any screening tool can be the trade-off between false positive and false negative results (or specificity and sensitivity), leading to unnecessary investigations in the former case and inefficiencies in the latter. An ideal test can be one with a high positive predictive value (PPV), minimizing unnecessary investigations, but detecting the vast majority of cancers. Another key factor can be the so-called "detection sensitivity" and also the lower limit of tumor size for detection, which is distinguished from test sensitivity. Unfortunately, waiting for the tumor to grow large enough to release circulating tumor markers at the necessary level for detection can be at odds with the requirement for early detection, which is to treat the tumor at a stage where treatment is most effective. Thus, there is a need for effective blood-based screening for early colorectal cancer based on circulating analytes.

[0088] Detection of circulating tumor DNA is increasingly being recognized as a viable "liquid biopsy" allowing for non-invasive detection and information investigation of tumors. In some cases, these techniques have been applied to colon, breast, and prostate cancer through identification of tumor-specific mutations. Sensitivity of these techniques can be limited due to the high background of normal (e.g., non-tumor-derived) DNA present in circulation.

[0089] Detection of tumor-specific methylation in blood can provide a distinct advantage over mutation detection. A number of single or multiple methylation biomarkers can be assessed in cancers including lung, colon, and breast cancer. These can have low sensitivity as they can not be sufficiently prevalent in tumors.

[0090] There remains a need for more sensitive and specific screening tools to detect early or low tumor burden colorectal cancer tumor signals in recurrence and for primary screening in high-risk populations.

[0091] The present disclosure provides methods and systems relating to genetic methylation profiling associated with colorectal cancer detection and disease progression.

[0092] On the one hand, this disclosure provides methods for using methylated region panels suitable for analyzing methylation within regions or genes; on the other hand, it provides new uses for said regions, genes, and gene products, as well as methods, assays, and kits relating to the detection, differentiation, and distinction of colonic proliferative disorders. The methods and nucleic acids provided herein can be used to analyze colonic proliferative disorders selected from adenocarcinoma, adenoma, polyp, squamous cell carcinoma, carcinoid tumor, sarcoma, and lymphoma.

[0093] In some embodiments, the method includes using one or more genes selected from methylated regions as biomarkers for the differentiation, detection, and distinction of colonic proliferative disorders. The use of these genes can be enabled by analyzing the methylation status of one or more genes selected from the methylated regions described herein, as well as their promoters or regulatory elements.

[0094] The methods and systems disclosed herein may include analyzing the methylation status of CpG dinucleotides in one or more genomic sequences based on the methylation regions and complementary sequences described herein.

[0095] I. Definition

[0096] Unless the context clearly indicates otherwise, as used in the specification and claims, the singular forms "a / an" and "the" include a plural of indicators. For example, the term "nucleic acid" includes a plurality of nucleic acids, including mixtures thereof.

[0097] As used herein, the term "object" generally refers to an entity or medium that has testable or detectable genetic information. An object can be a person, an individual, or a patient. An object can be a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, apes, farm animals, treadmills, rodents, and pets. An object can be a person who has cancer or is suspected of having cancer. An object can exhibit symptoms indicative of its health or physiological state or condition, such as cancer or other diseases, ailments, or symptoms. Alternatively, an object may be asymptomatic in relation to such health or physiological state or condition.

[0098] As used herein, the term "sample" generally refers to a biological sample obtained from or derived from one or more objects. A biological sample may be a cell-free biological sample or a substantially cell-free biological sample, or it may be processed or fractionated to produce a cell-free biological sample. For example, cell-free biological samples may include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and their derivatives. EDTA collection tubes and cell-free RNA collection tubes (e.g., [missing information]) may be used. ) or a cell-free DNA collection tube (e.g. ) from the subject. The cell-free biological sample can be derived from a whole blood sample by fractionation (e.g., centrifugation into a cellular component and a cell-free component). The biological sample or derivative thereof can contain cells. For example, the biological sample can be a blood sample or derivative thereof (e.g., blood collected by a collection tube or blood droplet).

[0099] As used herein, the term "nucleic acid" generally refers to a polymeric form of nucleotides of any length, whether deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nucleic acids can have any three-dimensional structure and can perform any function, known or unknown. Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coded or non-coded regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acid, branched nucleic acid, plasmid, vector, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleic acid can comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the nucleic acid. The sequence of nucleotides of a nucleic acid can be interrupted by non-nucleotide components. A nucleic acid can be further modified following assembly, such as by conjugation or association with a reporter.

[0100] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule in an initial population of nucleic acid molecules for which the presence, amount, and / or sequence of the nucleotide sequence of the nucleic acid molecule, or a change in one or more of these, needs to be determined. A target nucleic acid can be any type of nucleic acid, including DNA, RNA, and analogs thereof. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.

[0101] As used herein, the terms“amplifying” and“amplification” generally refer to increasing the size or number of nucleic acid molecules. The nucleic acid molecules can be single-stranded or double-stranded. Amplification can include generating one or more copies or“amplification products” of a nucleic acid molecule. Amplification can be performed, for example, by extension (e.g., primer extension) or ligation. Amplification can include performing a primer extension reaction to generate a strand complementary to a single-stranded nucleic acid molecule, and in some cases, to generate one or more copies of the strand and / or the single-stranded nucleic acid molecule. The term“DNA amplification” generally refers to generating one or more copies of a DNA molecule or“amplified DNA product.” The term“reverse transcription amplification” generally refers to generating deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template through the action of a reverse transcriptase enzyme

[0102] As used herein, the term“cell-free nucleic acid (cfNA)” generally refers to nucleic acids in a biological sample that are not contained within a cell, such as cell-free RNA (“cfRNA”) or cell-free DNA (“cfDNA”). cfDNA can circulate freely in bodily fluids, such as in the bloodstream.

[0103] As used herein, the term“cell-free sample” generally refers to a biological sample that is substantially devoid of intact cells. This can be derived from a biological sample that is itself substantially devoid of cells, or can be derived from a sample in which cells have been removed. Examples of cell-free samples include those derived from blood, such as serum or plasma; urine; or samples derived from other sources, such as semen, sputum, fecal matter, catheter effusions, lymph, or recovered lavage fluid.

[0104] As used herein, the term“circulating tumor DNA” generally refers to cfDNA derived from a tumor.

[0105] As used herein, the term“genomic region” generally refers to identified regions of nucleic acids that are identified according to their location in a chromosome. In some examples, a genomic region is referred to by a gene name and encompasses both coding and non-coding regions associated with a physical region of nucleic acid. As used herein, a gene includes coding regions (exons), non-coding regions (introns), transcriptional control regions or other regulatory regions, and promoters. In another example, a genomic region can incorporate an intron or exon or intron / exon boundary within a named gene.

[0106] As used herein, the term "CpG island" generally refers to a contiguous region of genomic DNA that satisfies the following criteria: (1) a frequency of CpG dinucleotides corresponding to an "observed / expected ratio" of greater than about 0.6; and (2) a "GC content" of greater than about 0.5. CpG islands are usually, but not always, between 0.2 and 3 kilobases (kb) in length, with a high frequency of CpG sites. CpG islands are found in the promoters or near the promoters of about 40% of mammalian genes. CpG islands are also found outside of mammalian genes. In some examples, CpG islands are found in exons, introns, promoters, enhancers, suppressors, and transcriptional regulatory elements. CpG islands can tend to occur upstream of so-called "housekeeping genes." CpG islands are said to have a CpG dinucleotide content of at least about 60% of the statistical expectation. The occurrence of CpG islands at or upstream of the 5' end of a gene can reflect a role in transcriptional regulation, and methylation of CpG sites within a gene's promoter can lead to silencing. Conversely, methylation-induced silencing of tumor suppressors is a hallmark of many human cancers.

[0107] As used herein, the term "CpG shore" generally refers to a short distance region extending outward from a CpG island, where methylation can also occur. CpG shores can be found in regions about 0 to 2 kb upstream and downstream of a CpG island.

[0108] As used herein, the term "CpG shelf" generally refers to a short distance region extending from a CpG shore, where methylation can also occur. CpG shelves can generally be found in regions between about 2 kb and 4 kb upstream and downstream of a CpG island (e.g., extending a further 2 kb outward from a CpG shore).

[0109] As used herein, the term "colonic cell proliferative disorder" generally refers to a disorder or disease that includes the disturbed or abnormal proliferation of a colon or rectal cell. In some examples, the disorder is selected from the group consisting of adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial cancer, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colonic cell proliferative disorder comprises colorectal cancer.

[0110] As used herein, the term "epigenetic parameter" generally refers to cytosine methylation. Further epigenetic parameters include, for example, acetylation of histones, although they can not be directly analyzed using the described methods, but rather, which are related to DNA methylation.

[0111] As used herein, the term "genetic parameter" generally refers to mutations and polymorphisms of genes and further desired sequences for their regulation. Examples of mutations include insertions, deletions, point mutations, inversions, and polymorphisms, such as SNPs (single nucleotide polymorphisms).

[0112] As used herein, the term "half-methylation" or "hemimethylation" generally refers to the methylation state of a palindromic CpG methylation site in which only one of the two cytosines in one of the CpG dinucleotide sequences of the palindromic CpG methylation site is methylated (e.g., 5'-CC M GG-3' (top strand): 3'-GGCC-5' (bottom strand).

[0113] As used herein, the term "hypermethylation" generally refers to the average methylation state corresponding to an increase in the presence of 5-mC at one or more CpG dinucleotides in the DNA sequence of a test DNA sample relative to the number of 5-mC found at the corresponding CpG dinucleotides in a normal control DNA sample. In some embodiments, the test DNA sample is from an individual having a colon cell proliferative disorder.

[0114] As used herein, the term "hypomethylation" generally refers to the average methylation state corresponding to a decrease in the presence of 5-mC at one or more CpG dinucleotides in the DNA sequence of a test DNA sample relative to the number of 5-mC found at the corresponding CpG dinucleotides in a normal control DNA sample. In some embodiments, the test DNA sample is from an individual having a colon cell proliferative disorder.

[0115] As used herein, the term "methylation state" or "methylation status" generally refers to the presence or absence of 5-methylcytosine ("5-mC") at one or more CpG dinucleotides in a DNA sequence. The methylation state of one or more specific CpG palindromic methylation sites (two CpG dinucleotide sequences per site) in a DNA sequence includes "unmethylated," "fully methylated," and "hemimethylated."

[0116] As used herein, the term "methylated cytosine" generally refers to any methylated form of the nucleic acid base cytosine, which contains a methyl or hydroxymethyl functional group at the 5' position. Methylated cytosines are known to be regulators of gene transcription in genomic DNA. This term can include 5-methylcytosine and 5-hydroxymethylcytosine.

[0117] As used herein, the term "methylation assay" generally refers to any assay for determining the methylation state of one or more CpG dinucleotide sequences within a DNA sequence.

[0118] As used herein, the term "minimal residual disease" or "MRD" generally refers to small amounts of cancer cells in the body after cancer treatment. MRD testing can be performed to determine whether cancer treatment is effective and to guide further treatment plans.

[0119] As used herein, the term "MSP" (methylation-specific polymerase chain reaction (PCR)) generally refers to a methylation assay, such as described by Herman Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996 and U.S. Patent No. 5,786,146, the contents of each of which are incorporated herein by reference.

[0120] As used herein, the term "methylation-converted" or "converted" nucleic acid generally refers to a nucleic acid, such as, for example, DNA, that has undergone a DNA conversion process for methylation sequencing. Examples of conversion processes include reagent-based (such as bisulfite) conversion, enzymatic conversion, or combined conversion (such as TET-assisted pyridine-borane sequencing (TAPS) conversion), in which unmethylated cytosines are converted to uracils prior to PCR amplification or sequencing. Conversion processes can be used in methylation sequencing methods to distinguish between methylated and unmethylated cytosine bases.

[0121] As used herein, the term "region methylated in cancer" generally refers to a segment of the genome that contains methylation sites (CpG dinucleotides) whose methylation is associated with a malignant cell state. The methylation of a region can be associated with more than one different type of cancer, or can be specifically associated with one type of cancer. Further, the methylation of a region can be associated with more than one cancer subtype, or can be specifically associated with one cancer subtype.

[0122] The terms "type" and "subtype" of cancer are used herein generally in a relative sense, whereby a "type" of cancer, such as breast cancer, can be a "subtype" based on, for example, stage, morphology, histology, gene expression, receptor profile, mutation profile, aggressiveness, prognosis, malignant characteristics, and the like. Likewise, "type" and "subtype" can be applied on a more granular level, for example, to distinguish a histological "type" into "subtypes," for example, defined according to mutation profile or gene expression. Cancer "stage" is also used to refer to a classification of cancer types based on histological and pathological characteristics associated with disease progression.

[0123] II. Analyzing a Sample

[0124] The cell-free biological sample can be obtained or derived from a human subject. The cell-free biological sample can be stored prior to processing under different storage conditions, such as different temperatures (e.g., room temperature, refrigerated or frozen conditions, 25°C, 4°C, -18°C, -20°C, or -80°C) or different suspensions (e.g., EDTA collection tubes, cell-free RNA collection tubes, or cell-free DNA collection tubes).

[0125] The cell-free biological sample can be obtained from a subject having cancer, a subject suspected of having cancer, or a subject not having or not suspected of having cancer.

[0126] Cell-free biological samples can be collected prior to and / or after treatment of a cancer subject. Cell-free biological samples can be obtained from a subject during treatment or a treatment regimen. Multiple cell-free biological samples can be obtained from a subject to monitor the effectiveness of treatment over time. Cell-free biological samples can be taken from a subject known or suspected to have cancer, while the subject cannot be definitively diagnosed as positive or negative by clinical testing. Samples can be taken from a subject suspected of having cancer. Cell-free biological samples can be taken from a subject presenting with unexplained symptoms such as fatigue, nausea, weight loss, pain, weakness, or bleeding. Cell-free biological samples can be taken from a subject with explained symptoms. Cell-free biological samples can be taken from a subject at risk of developing cancer due to factors such as family history, age, hypertension or pre-hypertension, diabetes or pre-diabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, drinking, or drug use), or presence of other risk factors.

[0127] Cell-free biological samples can comprise one or more analytes that can be analyzed, such as cell-free ribonucleic acid (cfRNA) molecules suitable for analysis to generate transcriptomic data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for analysis to generate genomic data, or mixtures or combinations thereof. One or more such analytes (e.g., cfRNA molecules and / or cfDNA molecules) can be isolated or extracted from one or more cell-free biological samples from a subject for downstream analysis using one or more suitable assays.

[0128] After a cell-free biological sample is obtained from a subject, the cell-free biological sample can be processed to generate a data set indicative of cancer in the subject. For example, nucleic acid molecules of the cell-free biological sample are assessed for presence, absence, or quantification (e.g., quantitative measures of RNA transcripts or DNA at cancer-related genomic loci) at a panel of loci of a cancer-related genome. In some embodiments, processing of a cell-free biological sample obtained from a subject can comprise: (i) placing the cell-free biological sample under conditions sufficient to isolate, enrich, or extract a plurality of nucleic acid molecules; and (ii) analyzing the plurality of nucleic acid molecules to generate a data set.

[0129] In some embodiments, a plurality of nucleic acid molecules is extracted from a cell-free biological sample and sequenced to generate a plurality of sequencing reads. Nucleic acid molecules can include ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). Nucleic acid molecules (e.g., RNA or DNA) can be extracted from a cell-free biological sample by a variety of methods, such as the QIAamp® Circulating Nucleic Acid Kit from QIAGEN®, the DNA Cell-Free Biological Mini Kit from QIAGEN®, or the Circulating Nucleic Acid Kit from Norgen Biomex® Kit Protocol, the DNA Cell-Free Biological Mini Kit from QIAGEN®, or the Circulating Nucleic Acid Kit from Norgen Biomex® Kit Protocol, the DNA Cell-Free Biological Mini Kit from QIAGEN®, or the Circulating Nucleic Acid Kit from Norgen ​​Qiagen® QIAamp® DNA Mini Kit protocol. The extraction method can extract all RNA or DNA molecules from the sample. Alternatively, the extraction method can selectively extract a portion of the RNA or DNA molecules from the sample. The RNA molecules extracted from the sample can be converted to DNA molecules by reverse transcription (RT).

[0130] Sequencing can be performed by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), ligation-based sequencing, hybridization sequencing, and

[0131] Sequencing can include nucleic acid amplification (e.g., of RNA or DNA molecules). In some embodiments, the nucleic acid amplification is polymerase chain reaction (PCR). Appropriate rounds of PCR (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to sufficiently amplify an initial amount of nucleic acid (e.g., RNA or DNA) to a desired input amount for subsequent sequencing. In some cases, PCR can be used for bulk amplification of target nucleic acids. This can include use of adaptor sequences that can be first ligated to different molecules, followed by PCR amplification using universal primers. PCR can be performed using any of a number of commercially available kits, such as the kits provided by Life Technologies® (Thermo Fisher Scientific®), Promega®, Bio-Rad®, and others. In other cases, only certain target nucleic acids within a population of nucleic acids can be amplified. Specific primers (possibly in combination with adaptor ligation) can be used to selectively amplify certain targets for downstream sequencing. PCR can include targeted amplification of one or more genomic loci, such as genomic loci associated with cancer. Sequencing can include use of simultaneous reverse transcription (RT) and polymerase chain reaction (PCR), such as the OneStep RT-PCR kit protocol provided by Life Technologies® (Thermo Fisher Scientific®). Thermo Fisher or OneStep RT-PCR kit protocol.

[0132] ​RNA or DNA molecules isolated or extracted from cell-free biological samples can be labeled, e.g., using identifiable tags, to allow multiplexing of multiple samples. Any number of RNA or DNA samples can be multiplexed. For example, a multiplexed reaction can comprise RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more than 100 initial cell-free biological samples. For example, multiple cell-free biological samples can be labeled with sample barcodes, such that each DNA molecule can be traced back to the sample (and subject) from which the DNA molecule originated. Such tags can be attached to RNA or DNA molecules by ligation or primer PCR amplification.

[0133] After sequencing the nucleic acid molecules, the sequence reads can be subjected to appropriate bioinformatic processing to generate data indicative of the presence, absence, or relative assessment of cancer. For example, the sequence reads can be aligned to one or more reference genomes (e.g., one or more species’ genomes, such as the human genome, e.g., hgl9). Aligned sequence reads can be quantified at one or more genomic loci to generate a data set indicative of cancer. For example, quantification of sequences corresponding to a plurality of genomic loci associated with cancer can generate a data set indicative of cancer.

[0134] Cell-free biological samples do not require any nucleic acid extraction to be processed. For example, a cancer in a subject can be identified or monitored by using probes configured to selectively enrich for nucleic acid (e.g., RNA or DNA) molecules corresponding to a plurality of cancer-associated genomic loci. The probes can be nucleic acid primers. The probes can have sequence complementarity to nucleic acid sequences from one or more of the plurality of cancer-associated genomic loci or genomic regions. The plurality of cancer-associated genomic loci or genomic regions can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, or more different cancer-associated genomic loci or genomic regions. The plurality of cancer-associated genomic loci or genomic regions can comprise one or more members selected from the groups listed in Tables 1-11 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, or more). The cancer-associated genomic loci or genomic regions can be associated with different stages or subtypes of cancer (e.g., colorectal cancer).

[0135] The probes can be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity to nucleic acid sequences (e.g., RNA or DNA) of one or more genomic loci (e.g., cancer-associated genomic loci). These nucleic acid molecules can be primers or enrichment sequences. Analysis of a cell-free biological sample using probes selective for one or more genomic loci (e.g., cancer-associated genomic loci) can include using array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some embodiments, DNA or RNA can be analyzed by one or more of the following: DNA / RNA isothermal amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassay, electrochemical assay, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assay, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (e.g., CRISPR typing PCR (ctPCR), specific high-sensitivity enzymatic reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR trans reporter (DETECTR), and CRISPR-mediated analog multi-event recording device (CAMERA)), and laser transmission spectroscopy (LTS).

[0136] The assay readout can be quantified on one or more genomic loci (e.g., cancer-associated genomic loci) to generate data indicative of cancer. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to a plurality of genomic loci (e.g., cancer-associated genomic loci) can generate data indicative of cancer. The assay readout can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, and the like, or normalized values thereof. The assay can be a home user test configured to be performed in a home environment.

[0137] In some embodiments, multiplexed assays can be used to simultaneously process cell-free biological samples of a subject. For example, a first assay can be used to process a first cell-free biological sample obtained or derived from a subject to generate a first data set indicative of cancer; and a second assay, different from the first assay, can be used to process a second cell-free biological sample obtained or derived from the subject to generate a second data set indicative of cancer. Either or both of the first data set and the second data set can then be analyzed to assess cancer of the subject. For example, a single diagnostic indicator or diagnostic score can be generated based on a combination of the first data set and the second data set. Another example is that separate diagnostic indicators or diagnostic scores can be generated from the first data set and the second data set.

[0138] Cell-free biological samples can be processed using methylation-specific assays. For example, methylation-specific assays can be used to identify a quantitative measure of methylation (e.g., indicative of presence, absence, or relative quantity) of each of a plurality of cancer-associated genomic loci in a cell-free biological sample of a subject. Methylation-specific assays can be configured to process cell-free biological samples, such as a blood sample or a urine sample (or derivative thereof) of a subject. A quantitative measure of methylation (e.g., indicative of presence, absence, or relative quantity) of cancer-associated genomic loci in a cell-free biological sample can be indicative of one or more cancers. Methylation-specific assays can be used to generate a dataset indicative of a quantitative measure of methylation (e.g., indicative of presence, absence, or relative quantity) of each of a plurality of cancer-associated genomic loci in a cell-free biological sample of a subject.

[0139] For example, methylation-specific assays can include one or more of methylation- aware sequencing (e.g., using bisulfite treatment), pyrosequencing, methylation-sensitive single- strand conformation analysis (MS-SSCA), high-resolution melting analysis (HRM), methylation-sensitive single-nucleotide primer extension (MS-SnuPE), base-specific cleavage / MALDI-TOF, microarray-based methylation assays, methylation-specific PCR, targeted bisulfite sequencing, oxidative bisulfite sequencing, mass spectrometry-based bisulfite sequencing, or reduced representation bisulfite sequencing (RRBS).

[0140] III. Signature Panel

[0141] The present disclosure provides methods and systems for analyzing a biological sample to obtain measurable features from a combination of hypermethylated regions of DNA in the sample that are associated with the development of a colon cell proliferative disorder, thereby identifying a signature panel of regions. Features from the signature panel can be processed using a trained algorithm (e.g., a machine learning model) to create a classifier configured for stratifying a population of individuals for a colon cell proliferative disorder. The methods feature the use of one or more nucleic acids having a methylation region described in the signature panel that are contacted with one or a series of reagents capable of distinguishing between methylated and non-methylated CpG dinucleotides within the identified region prior to sequencing.

[0142] The signature panel described herein generally refers to a collection of genomic DNA targeted regions identified in a cell-free nucleic acid sample and exhibiting increased cytosine base methylation in samples associated with a colon cell proliferative disorder. The formation of the signature panel allows for rapid and specific analysis of particular methylation regions associated with a colon cell proliferative disorder. The signature panel described and used in the methods herein can be used to improve diagnosis, prognosis, treatment selection, and monitoring (e.g., treatment monitoring) of a colon cell proliferative disorder.

[0143] The signature panels and methods of the present disclosure can provide a significant improvement over current methods in addressing the need for markers or signature panels for detecting early colon cell proliferative disorders from bodily fluid samples such as whole blood, plasma, or serum. Current methods for detecting and diagnosing colon cell proliferative disorders include colonoscopy, sigmoidoscopy, and fecal occult blood colon cancer. In comparison to these methods, the methods provided herein can be much less invasive than colonoscopy and at least as sensitive or more sensitive than sigmoidoscopy, fecal immunochemical test (FIT), and fecal occult blood test (FOBT). In comparison to these markers currently in use, the methods provided herein can have a significant advantage in sensitivity and specificity due to the use of a gene panel in advantageous combination with a high sensitivity assay technique.

[0144] In some embodiments, the region methylated in the cancer comprises a CpG island. In some embodiments, the region methylated in the cancer comprises a CpG shore. In some embodiments, the region methylated in the cancer comprises a CpG shelf. In some embodiments, the region methylated in the cancer comprises a CpG island and a CpG shore. In some embodiments, the region methylated in the cancer comprises a CpG island, a CpG shore, and a CpG shelf.

[0145] In some embodiments, the region methylated in the cancer comprises a CpG island and sequences upstream and downstream of about 0 to 4 kilobases (kb). The region methylated in the cancer can also comprise a CpG island and sequences upstream and downstream of about 0 to 3 kb, about 0 to 2 kb, about 0 to 1 kb, about 0 to 500 base pairs (bp), about 0 to 400 bp, about 0 to 300 bp, about 0 to 200 bp, or about 0 to 100 bp.

[0146] According to some examples, a number of design parameters can be considered in selecting a hypermethylated region in a cancer. In certain examples, the length of the methylated region is about 200 bp, about 300 bp, about 400 bp, or about 500 bp. Data for this selection process can be obtained from a variety of sources, such as, for example, The Cancer Genome Atlas (TCGA) (cancergenome.nih.gov) by using, for example, the Cancer Genome Atlas (TCGA) database for a wide variety of cancers. Infinium HumanMethylation450 BeadChip, or from other sources based on bisulfite whole genome sequencing or other methods. In some embodiments, a region can be selected using "methylation values" (which can be derived from TCGA level 3 methylation data, which in turn is derived from beta values of about -0.5 to 0.5) In some embodiments, a primer set is used for amplification that is designed to amplify at least one methylation site that has a methylation value that is lower than about -0.3 in normal cases. This can be established in a plurality of normal tissue samples, such as about 4. The methylation value can be equal to or lower than about -0.1, about -0.2, about -0.3, about -0.4, about -0.5, about -0.6, about -0.7, about -0.8, about -0.9, or about -1.0.

[0147] In some embodiments, a primer set is designed to amplify at least one methylation site that has a difference between the average methylation values in cancerous tissue and normal tissue that is greater than a predefined threshold, such as about 0.3. In some embodiments, the difference can be greater than about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0. In some examples, the proximity of other methylation sites that satisfy this requirement can also play a role in selecting a region. In some embodiments, a primer set includes a primer pair that amplifies at least one methylation site that has at least one methylation site within about 200 bp, has a methylation value that is also lower than about -0.3 in normal cases, and has a difference between the average methylation values in cancerous tissue and normal tissue that is greater than about 0.3.

[0148] In some examples, a target region is selected if a region is more methylated than the same region in a sample obtained or derived from one or more healthy individuals (e.g., individuals without cancer). This selection can be performed manually or in a computational manner. In certain examples, a region is selected if it is more methylated by at least about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, or more than about 100% than a sample from a healthy individual. In another example, a region can be selected if the number of reads mapping to the region in a disease sample at a pre-defined threshold methylation CpG count exceeds the same pre-defined threshold methylation CpG count for the same region in a healthy individual sample. For a given region, the methylation CpG count used as a baseline threshold in a healthy sample can vary, but the number of reads mapping to the region exceeding the baseline threshold of the methylation CpG count for the region in a healthy sample can indicate an important region regardless of fluctuations in the threshold CpG count.

[0149] In some examples, a target region can be selected for amplification based on the number of samples in the validation set that have methylation at the site. For example, a region can be selected if at least about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% of the samples tested from disease individuals have a higher degree of methylation compared to a sample from a healthy individual. For example, regions can be selected if they are methylated in at least about 75% of the tested tumors (including within a particular subtype). For some validations, cell lines derived from tumors can be used for testing.

[0150] The present disclosure also provides a method for performing an assay to determine genetic and / or epigenetic parameters of one or more genes selected from the signature panels described herein and their promoters and regulatory elements. In some embodiments, the assay is performed according to a method to detect methylation within one or more genes selected from the signature panels described herein, wherein the methylated nucleic acids are present in a solution that also contains excess background DNA, wherein the background DNA is present at about 100 to 1000 fold, about 100 to 10,000 fold, about 100 to 100,000 fold, about 1000 to 10,000 fold, about 1000 to 100,000 fold, or about 10,000 to 100,000 fold of the concentration of DNA to be detected. In some embodiments, the concentration of DNA to be detected is greater than about 100,000 fold of the concentration of background DNA. In some embodiments, the method comprises contacting a nucleic acid sample obtained from a subject with at least one reagent or series of reagents (e.g., a reagent that distinguishes between methylated and non-methylated CpG dinucleotides within a target nucleic acid).

[0151] The tumor or colon cell proliferative disorder as described herein can be selected from the group consisting of: adenoma (adenomatous polyps), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial cancer, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal cancer.

[0152] The signature panel comprising information methylation regions can be selected according to the purpose of the intended assay. For targeted approaches, primer pairs can be designed based on the intended set of target regions. In some embodiments, the set of regions comprises at least one, at least two, at least three, or more than three of the regions listed in Table 1. In some embodiments, the set of regions comprises all of the regions listed in Table 1.

[0153] In some embodiments, the set of methylation regions associated with colorectal cancer is selected from Table 1.

[0154] In some embodiments, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, SFMBT2, CHST10, CCNA1, BEND4, KRBA1, S1PR1, PPP1R16B, IKZF1, LONRF2, ZFP82, and FLT3 (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises all of the regions listed in Table 1. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, SFMBT2, CHST10, CCNA1, BEND4, KRBA1, S1PR1, PPP1R16B, IKZF1, LONRF2, ZFP82, and FLT3.

[0155] Table 1

[0156]

[0157]

[0158] In some embodiments, the method further comprises quantifying the methylation signal, wherein a value that exceeds a predetermined threshold indicates a colon cell proliferative disorder. In some embodiments, the quantification and comparison of each methylation site in a colon cell proliferative disorder is performed independently. Thus, a count of tumor positive signals can be established for each site. In some embodiments, the method further comprises determining the proportion of sequencing reads that comprise a tumor signal, wherein a proportion that exceeds a threshold indicates a colon cell proliferative disorder. In some embodiments, the determination of each methylation site in a colon cell proliferative disorder is performed independently.

[0159] As used herein, the term "threshold" generally refers to a value selected to distinguish, separate, or differentiate between two populations of objects. In some embodiments, a threshold value distinguishes between a disease (e.g., malignant) state and a non-disease (e.g., healthy) state. In some embodiments, a threshold value can distinguish between different stages of a disease (e.g., stage 1, stage 2, stage 3, or stage 4). A threshold value can be set according to the disease in question, and can be determined according to an early analysis, such as an analysis of a training set, or according to a set of inputs with known characteristics (e.g., healthy, disease, or disease stage). Threshold values can also be set for a gene region based on the methylation prediction values for particular sites. Threshold values can differ for each methylation site, and data from multiple sites can be combined in a final analysis.

[0160] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from: ITGA4, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 2. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from: ITGA4, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.

[0161] Table 2

[0162] Methyl regions (Gene ID; Chromosome: position start - position end) ITGA4; chr2: 181457004-181457950 TMEM163; chr2: 134718243-134719428 SFMBT2; chr10: 7408046-7408953 ELMO1; chr7: 37448612-37449471 ZNF543; chr19: 57320164-57320845 SFMBT2; chr10: 7410025-7411008 CHST10; chr2: 100417269-100417795 ELMO1; chr7: 37447852-37448217 CCNA1; chr13: 36431498-36432414 BEND4; chr4: 42150707-42153216 KRBA1; chr7: 149714695-149715338 S1PR1; chr1: 101236505-101237190 PPP1R16B; chr20: 38805341-38807221

[0163] In some embodiments, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from: EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 3. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from: EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.

[0164] Table 3

[0165] Methyl regions (Gene ID; Chromosome: position start - position end) EMBP1; chr1: 121519076-121519744 TMEM163; chr2: 134718243-134719428 SFMBT2; chr10: 7408046-7408953 ELMO1; chr7: 37448612-37449471 ZNF543; chr19: 57320164-57320845 SFMBT2; chr10: 7410025-7411008 CHST10; chr2: 100417269-100417795 ELMO1; chr7: 37447852-37448217 CCNA1; chr13: 36431498-36432414 BEND4; chr4: 42150707-42153216 KRBA1; chr7: 149714695-149715338 S1PR1; chr1: 101236505-101237190 PPP1R16B; chr20: 38805341-38807221

[0166] In some embodiments, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, and S1PR1, and the tumor is a colorectal cancer. In some embodiments, the cancer panel comprises one or more regions listed in Table 4. In some embodiments, the probe is directed to a sequence selected from at least one, at least two, at least three, or more than three of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, and S1PR1.

[0167] Table 4

[0168] Methyl regions (Gene ID; Chromosome: position start - position end) ITGA4; chr2: 181457004-181457950 EMBP1; chr1: 121519076-121519744 TMEM163; chr2: 134718243-134719428 SFMBT2; chr10: 7408046-7408953 ELMO1; chr7: 37448612-37449471 ZNF543; chr19: 57320164-57320845 SFMBT2; chr10: 7410025-7411008 CHST10; chr2: 100417269-100417795 ELMO1; chr7: 37447852-37448217 CCNA1; chr13: 36431498-36432414 BEND4; chr4: 42150707-42153216 KRBA1; chr7: 149714695-149715338 S1PR1; chr1: 101236505-101237190

[0169] In some embodiments, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, and ZNF543, and the tumor is a colorectal cancer. In some embodiments, the cancer panel comprises a region listed in Table 5. In some embodiments, the probe is directed to a sequence selected from at least one, at least two, at least three, or more than three of: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, and ZNF5431.

[0170] Table 5

[0171] Methyl regions (Gene ID; Chromosome: position start - position end) ITGA4; chr2: 181457004-181457950 EMBP1; chr1: 121519076-121519744 TMEM163; chr2: 134718243-134719428 SFMBT2; chr10: 7408046-7408953 ELMO1; chr7: 37448612-37449471 ZNF543; chr19: 57320164-57320845

[0172] In some embodiments, the cancer panel comprises one or more of regions ITGA4 and EMBP1 (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 6. In some embodiments, the probe is directed to a sequence comprising ITGA4 and EMBP1.

[0173] Table 6

[0174] Methyl regions (Gene ID; Chromosome: position start - position end) ITGA4; chr2: 181457004-181457950 EMBP1; chr1: 121519076-121519744

[0175] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from KZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, IRF4, LOC100130992, CHST11, CHST15, RASSF2, EMILIN2, TMEM163, CHST10, and HCK (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 7. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, IRF4, LOC100130992, CHST11, CHST15, RASSF2, EMILIN2, TMEM163, CHST10, and HCK.

[0176] Table 7

[0177]

[0178]

[0179] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, and IRF4 (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 8. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, and IRF4.

[0180] Table 8

[0181]

[0182]

[0183] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1 (e.g., where the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 9. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1.

[0184] Table 9

[0185]

[0186]

[0187] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1 (e.g., wherein the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 10. In some embodiments, the probes are directed to at least one, at least two, at least three, or more than three sequences selected from IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1.

[0188] Table 10

[0189] Methyl regions (Gene ID; Chromosome: position start - position end) IKZF1; chr7: 50303445-50305526 KCNQ5; chr6: 72620772-72623556 ELMO1; chr7: 37447220-37450201 CHST2; chr3: 143118680-143121423 PRKCB; chr16: 23835445-23837405 FLI1; chr11: 128691887-128696541

[0190] In some embodiments of the above methods, the cancer panel comprises at least one, at least two, or at least three regions selected from IKZF1, KCNQ5, and ELMO1 (e.g., wherein the tumor is a colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 11. In some embodiments, the probes are directed to at least one, at least two, or at least three sequences selected from IKZF1, KCNQ5, and ELMO1.

[0191] Table 11

[0192] Methyl regions (Gene ID; Chromosome: position start - position end) IKZF1; chr7: 50303445-50305526 KCNQ5; chr6: 72620772-72623556 ELMO1; chr7: 37447220-37450201

[0193] In one aspect, the present disclosure provides a method for identifying a methylation signature indicative of a biological characteristic, the method comprising: obtaining, for a population, data comprising a plurality of genomic methylation datasets associated with a colon cell proliferative disorder status, each of the genomic methylation datasets being associated with biological information of a corresponding sample; separating the methylation datasets into a first group corresponding to a tissue or cell type having the biological characteristic and a second group corresponding to a plurality of tissues or cell types not having the biological characteristic; matching the methylation data of the first group with the methylation data of the second group on a site-by-site basis across the genome; identifying, on a site-by-site basis across the genome, a set of CpG sites that satisfy a predetermined threshold for establishing differential methylation between the first group and the second group; using the set of CpG sites to identify a target genomic region comprising at least one, at least two, at least three, or more than three differentially methylated CpGs within about 30 to 300 bp that satisfy the predetermined criteria to identify a differentially methylated genomic region, thereby providing a methylation signature indicative of a biological characteristic associated with the presence of a colon cell proliferative disorder.

[0194] In some instances, the target genomic region contains at least one, at least two, at least three, or more than three differentially methylated CpG sites within regions of the following lengths: approximately 30 to 150 bp, approximately 40 to 150 bp, approximately 50 to 150 bp, approximately 75 to 150 bp, approximately 100 to 150 bp, approximately 150 to 300 bp, approximately 150 to 250 bp, approximately 150 to 200 bp, approximately 200 to 300 bp, or approximately 250 to 300 bp.

[0195] In some instances, the target genomic region contains at least four differentially methylated CpG sites, at least four differentially methylated CpG sites, at least five differentially methylated CpG sites, at least six differentially methylated CpG sites, at least seven differentially methylated CpG sites, at least eight differentially methylated CpG sites, at least nine differentially methylated CpG sites, at least ten differentially methylated CpG sites, at least 12 differentially methylated CpG sites, or at least 15 differentially methylated CpG sites.

[0196] In some embodiments, the method further includes verifying the extended target genome region by detecting differential methylation within the extended target genome region using DNA from at least one independent sample possessing the biological trait and DNA from at least one independent sample not possessing the biological sample.

[0197] In some embodiments, the identification further includes limiting the CpG site set to CpG sites that further exhibit differential methylation compared to peripheral blood mononuclear cells from reference or control samples.

[0198] In some implementations, a predetermined threshold of at least about 50% methylation is set in the first group.

[0199] In some implementations, the predetermined threshold is the average methylation difference between the first group and the second group of at least about 0.3.

[0200] In some implementations, biological characteristics include malignancy.

[0201] In some implementations, biological traits include cancer type.

[0202] In some implementations, biological traits include cancer stage.

[0203] In some implementations, biological traits include cancer classification.

[0204] In some implementation schemes, cancer classification includes cancer grading.

[0205] In some implementation schemes, cancer classification includes histological classification.

[0206] In some embodiments, the biological trait comprises a metabolic profile.

[0207] In some embodiments, the biological trait comprises a mutation.

[0208] In some embodiments, the mutation is a disease-associated mutation.

[0209] In some embodiments, the biological trait comprises a clinical outcome.

[0210] In some embodiments, the biological trait comprises a drug response.

[0211] In some embodiments, the method further comprises designing a plurality of PCR primer pairs to amplify portions of the extended target genomic region, each portion comprising at least one differentially methylated CpG site.

[0212] In some embodiments, the design of the plurality of primer pairs comprises converting unmethylated cytosines to uracils to mimic conversion of cytosines to uracils, and designing primer pairs using the converted sequences.

[0213] In some embodiments, the primer pairs are designed to have a methylation bias.

[0214] In some embodiments, the primer pairs are methylation specific.

[0215] In some embodiments, there are no CpG residues within the primer pairs, and no bias towards methylation status.

[0216] In one aspect, the disclosure provides a method for synthesizing primer pairs specific to a methylation signature, the method comprising: performing the method of the disclosure and synthesizing the designed primer pairs.

[0217] IV. Nucleic acid conversion and methylation sequencing

[0218] A. Nucleic acid processing

[0219] Methylation sequencing can utilize a variety of methods, including chemical and enzymatic conversion of nucleic acid bases to distinguish between methylated cytosine and unmethylated cytosine in a nucleic acid sequence. These assays allow determination of the methylation status of one or more CpG dinucleotides (e.g., CpG islands) within a DNA sequence. Such assays can include, among others, bisulfite treatment of DNA sequencing or enzymatic treatment of DNA sequencing, polymerase chain reaction (PCR) (for sequence-specific amplification), quantitative PCR (qPCR), or digital droplet PCR (ddPCR), Southern blot analysis. In various examples, DNA in a biological sample is treated in such a way that cytosine bases unmethylated at the 5'-position are converted to uracil, thymine, or another base that differs in hybridization behavior from cytosine. This can be referred to as "conversion."

[0220] In some embodiments, the reagent converts cytosine bases unmethylated at the 5'-position to uracil, thymine, or another base that differs in hybridization behavior from cytosine.

[0221] Bisulfite modification of DNA generally refers to a tool for assessing CpG methylation status. A common method for analyzing the presence of 5-methylcytosine (5-mC) in DNA is based on the reaction of bisulfite with cytosine, which is converted to uracil under subsequent alkaline desulfurization, corresponding to the base pairing behavior of thymine. For example, by using bisulfite treatment, genomic sequencing has been adapted for the analysis of DNA methylation patterns and 5-methylcytosine distribution (e.g., as described by Frommer et al., Proc. Natl. Acad. Sci. USA 89: 1827-1831, 1992, the contents of which are incorporated herein by reference). Notably, however, 5-methylcytosine remains unmodified under these conditions. Thus, the original DNA is converted in such a way that methylcytosine (methyl-C), which initially cannot be distinguished from cytosine by hybridization behavior, can now be detected as the only remaining cytosine by various molecular biology techniques, for example, by amplification and hybridization or by sequencing. In various examples, other reagents can achieve the same results as bisulfite modification suitable for methylation sequencing.

[0222] A commonly used direct sequencing method employs PCR-amplified bisulfite-treated DNA, which is suitable for whole genome bisulfite sequencing (WGBS) or targeted bisulfite sequencing.

[0223] Bisulfite sequencing can refer to a commercially available NGS method for assessing site-specific DNA methylation changes. Probes are designed to be strand-specific and bisulfite-specific. Both methylated and non-methylated sequences are amplified. The process is similar to pyrosequencing, but overall provides higher throughput. In some embodiments, next-generation sequencing platforms are used to deliver large amounts of useful DNA methylation information (e.g., EPIGENTEK, Farmingdale, NY and ZYMO RESEARCH, Irvine, CA). By subjecting DNA to bisulfite treatment, then PCR amplification of target regions, library construction, and sequencing of amplicon regions, single-base resolution methylation analysis of individual cytosines in DNA can be facilitated. Specific primers can be designed for regions of interest, and changes in cytosine methylation within that region can be assessed. Each target DNA methylation site can be assessed at high sequencing coverage depth to obtain accurate, quantitative, and single-base resolution data output.

[0224] Enzymatic Methyl-seq (EM-seq) can rely on enzymatic conversion of nucleic acids for methylome analysis. Data can suggest that the process of generating EM-seq libraries does not destroy DNA as much as bisulfite sequencing. EM-seq libraries, while using fewer PCR cycles for all DNA input amounts, can have higher PCR yield, which suggests less DNA is lost during enzymatic treatment and library preparation compared to whole genome bisulfite sequencing (WGBS). In turn, the reduced PCR cycles can translate into more complex libraries and fewer PCR replicates during sequencing. The average insert size of EM-seq libraries can also be larger than WGBS, further supporting the fact that DNA remains intact. In the EM-seq workflow, TET2 oxidizes 5-mC and 5-hmC, which is protected from APOBEC deamination in the next operation. In contrast, unmodified cytosines are deaminated to uracils. In some embodiments, targeted methods include enzymatic conversion of nucleic acids (TEM-seq). In some embodiments, methylation sequencing methods are accomplished with Enzymatic Methyl-seq (New England Biolabs, Ipswich, MA), which is useful for the identification of 5mC and 5hmC. Enzymatic Methyl-seq (EM-seq) can rely on enzymatic conversion of nucleic acids for methylome analysis. Data can suggest that the process of generating EM-seq libraries does not destroy DNA as much as bisulfite sequencing. EM-seq libraries, while using fewer PCR cycles for all DNA input amounts, can have higher PCR yield, which suggests less DNA is lost during enzymatic treatment and library preparation compared to whole genome bisulfite sequencing (WGBS). In turn, the reduced PCR cycles can translate into more complex libraries and fewer PCR replicates during sequencing. The average insert size of EM-seq libraries can also be larger than WGBS, further supporting the fact that DNA remains intact. In the EM-seq workflow, TET2 oxidizes 5-mC and 5-hmC, which is protected from APOBEC deamination in the next operation. In contrast, unmodified cytosines are deaminated to uracils. In some embodiments, targeted methods include enzymatic conversion of nucleic acids (TEM-seq). In some embodiments, methylation sequencing methods are accomplished with Enzymatic Methyl-seq (New England Biolabs, Ipswich, MA), which is useful for the identification of 5mC and 5hmC.

[0225] In another example, 5hmC can also be sequenced using TET-assisted bisulfite sequencing (TAB-seq) (e.g., as described by Yu, M. et al. (2012). Nat. Protoc. 7, 2159-2170, the contents of which are incorporated herein by reference) (WiseGene; Fragmented DNA can be enzymatically modified using continuous T4 bacteriophage beta-glucosyltransferase (T4-BGT) followed by treatment with 10-11 transposase (TET) dioxygenase prior to bisulfite addition. T4-BGT glycosylates 5hmC to form beta-glucosyl-5-hydroxymethylcytosine (5ghmC) which is then oxidized to 5caC with TET. Only 5ghmC is not subject to subsequent deamination by bisulfite, which allows 5hmC to be distinguished from 5mC by sequencing.

[0226] Oxidative bisulfite sequencing (oxBS) provides another method to distinguish 5mC from 5hmC (e.g., as described by Booth, M.J., et al., 2012 Science 336:934-937, the contents of which are incorporated herein by reference). The oxidizing reagent potassium peroxymonocarbonate converts 5hmC to 5-formylcytosine (5fC), which is deaminated to generate uracil by subsequent bisulfite treatment. 5mC remains unchanged and can therefore be identified using this method.

[0227] APOBEC-coupled epigenetic sequencing (ACE-seq) completely excludes bisulfite conversion and relies on enzymatic conversion to detect 5hmC (e.g., as described by Schutsky, E.K. et al., Nat. Biotechnol., 2018 Oct 8, the contents of which are incorporated herein by reference). With this method, T4-BGT glycosylates 5hmC to 5ghmC and protects it from deamination by apolipoprotein B mRNA-editing enzyme catalytic subunit 3A (APOBEC3A). Cytosine and 5mC are deaminated by APOBEC3A and sequenced as thymine.

[0228] In another example, a bisulfite-free and base-level resolution sequencing method, TET-assisted pyridine-borane sequencing (TAPS), can be used for detection of 5mC and 5hmC. TAPS combines 10-11 transposase (TET) oxidation of 5mC and 5hmC to 5-carboxylcytosine (5caC) with pyridine-borane reduction of 5caC to dihydrouracil (DHU). Subsequent PCR converts DHU to thymine, enabling C-to-T conversion of 5mC and 5hmC. TAPS directly detects modifications with high sensitivity and specificity without affecting unmodified cytosine. (e.g., as described by Liu, Y. et al. Nat Biotechnol. 2019 Apr; 37(4): 424-429, the contents of which are incorporated herein by reference).

[0229] TET-assisted 5-methylcytosine sequencing (TAmC-seq) enriches 5mC loci and utilizes two sequential enzymatic reactions followed by affinity pull-down (e.g., as described by Zhang, L. 2013, Nat Commun 4: 1517, the contents of which are incorporated herein by reference). Fragmented DNA is treated with T4-BGT to protect 5hmC by glycosylation. The mTET1 enzyme is then used to oxidize 5mC to 5hmC, and the newly formed 5hmC is labeled with T4-BGT using a modified glucose moiety (6-N3-glucose). Click chemistry is used to introduce a biotin tag, enabling enrichment of 5mc-containing DNA fragments for detection and whole-genome profiling.

[0230] B. Next-Generation Sequencing

[0231] In some embodiments, the generation of sequencing reads is performed by next- generation sequencing. This can allow for higher read depth to be achieved for a given region. These can be high-throughput methods, including, for example (Illumina) sequencing, DNB-Sequencer T7 or G400 (MGI Tech Co., Ltd), sequencing (GenapSys, Inc.), Roche 454 sequencing (Roche Sequencing Solutions, Inc.), Ion Torrent sequencing (Thermo Fisher Scientific), and SOLiD sequencing (Thermo Fisher ) sequencing. The number of sequencing reads can be adjusted according to the amount of DNA input and the depth of data desired for analysis.

[0232] In some embodiments, the generation of sequencing reads is performed simultaneously on samples obtained from multiple patients, with each patient’s cell-free nucleic acid fragments being barcoded. This allows for parallel analysis of multiple patients in one sequencing run.

[0233] In another aspect, the disclosure provides a kit for detecting a tumor, comprising reagents for performing the above-described methods and instructions for detecting a tumor signal. The reagents can include, for example, a primer set, PCR reaction components, and / or sequencing reagents.

[0234] C. Targeted Sequencing

[0235] In targeted methylation sequencing methods, targeted regions in a biological sample, such as cfDNA, are analyzed in order to determine the methylation status of the targeted gene sequences. In some embodiments, the targeted regions include adjacent nucleotides of a target region of interest, such as at least about 16 adjacent nucleotides of a target region of interest, or hybridize thereto under stringent conditions. In different examples, targeted sequencing can be achieved using hybrid capture and amplicon sequencing methods.

[0236] D. Hybrid Capture

[0237] The hybridization methods provided herein can be used for various forms of nucleic acid hybridization, such as in-solution hybridization and hybridization such as on solid supports (e.g., membranes, microarrays, and RNA, DNA, and in situ hybridization on cell / tissue slides). In particular, the methods are suitable for in-solution hybridization capture for target enrichment of certain types of genomic DNA sequences (e.g., exons) used in next generation targeted sequencing. For the hybrid capture method, cell-free nucleic acid samples undergo library preparation. As used herein, “library preparation” includes end repair, A-tailing, adapter ligation, or any other preparation performed on cell-free DNA to allow for subsequent DNA sequencing. In certain examples, the prepared cell-free nucleic acid library sequences contain adapters, sequence tags, index barcodes attached to the cell-free nucleic acid sample molecules. Various commercially available kits can be utilized to aid in library preparation for next generation sequencing methods. The construction of next generation sequencing libraries can include the use of a coordinated series of enzymatic reactions to prepare nucleic acid targets to generate a collection of random DNA fragments of a specific size for high-throughput sequencing. Advances and developments in various library preparation techniques have expanded the use of next generation sequencing in areas such as transcriptomics and epigenetics.

[0238] Improvements in sequencing technology have brought about changes and improvements in library preparation. Next generation sequencing library preparation kits developed by companies such as Bioo Kapa New England Life Pacific and have provided consistency and reproducibility for various molecular biology reactions, ensuring compatibility with the latest NGS instrument technology.

[0239] In different examples of targeted capture gene panels, various library preparation kits can be selected from Nextera Flex DNA Prep Ion (ThermoFisher ), (Thermo Fisher Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl Agilent SureSelect Human Methyl

[0240] In some embodiments, a hybrid capture method is performed on the prepared library sequence using specific probes. In some embodiments, the term "specific probe" as used herein generally refers to a probe specific to a known methylation site. In some embodiments, the design of the specific probe is based on using the human genome as a reference sequence and using a specific genomic region known to have a methylation site as a target sequence. Specifically, the genomic region known to have a methylation site can include at least one of the following regions: a promoter region, a CpG island region, a CGI island shore region, and a imprinting gene region. Thus, when hybrid capture is performed using the specific probe of some embodiments, sequences in the sample genome complementary to the target sequence, e.g., regions in the sample genome known to have a methylation site (also referred to herein as "specific genomic regions"), can be effectively captured.

[0241] According to one example, the methylation regions described herein are used to design specific probes. In some embodiments, the specific probes are designed using a commercially available method, such as, for example, the eArray system. The length of the probes can be sufficient to hybridize to the target methylation region with sufficient specificity. In various examples, the probes are 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, or 20-mer.

[0242] The regions listed in Tables 1-11 above are screened using database resources, such as Gene Ontology. According to the principle of complementary base pairing, a single-stranded capture probe can be combined complementarily with a single-stranded target sequence, thereby successfully capturing the target region. In some embodiments, the designed probes can be designed as a solid capture chip (in which the probes are fixed on a solid support) or as a liquid capture chip (in which the probes are free in a liquid), but are limited by various factors, such as probe length, probe density, and high cost, etc., and the solid capture chip is rarely used, while the liquid capture chip is used more.

[0243] In some embodiments, GC-rich sequences in the nucleic acid (where GC base content is higher than 60%) can result in reduced capture efficiency due to the molecular structure of C and G bases, as compared to normal sequences (where A, T, C, G base average content is 25% respectively). For areas of interest, such as CGI regions (CpG islands), it can be advisable to design a larger number of probes to obtain sufficient and accurate CGI data.

[0244] E. Sequencing based on amplicons

[0245] The converted DNA fragments can be amplified. In some embodiments, amplification is performed with primers designed to anneal to the methylation converted target sequence with at least one methylation site therein. Methylation sequencing conversion results in unmethylated cytosines being converted to uracils, while 5-methylcytosines are unaffected. Thus, a "converted target sequence" is understood to be a sequence in which cytosines known to be methylation sites are fixed as "C" (cytosine), while cytosines known to be unmethylated are fixed as "U" (uracil; which can be treated as "T" (thymine) in primer design).

[0246] In various examples, the source of DNA is cell-free DNA from whole blood, plasma, serum, or genomic DNA extracted from cells or tissue. In some embodiments, the amplified fragments are between about 100 and 200 base pairs in length. In some embodiments, the DNA source is extracted from a cell source (e.g., tissue, biopsy, cell line), and the amplified fragments are between about 100 and 350 base pairs in length. In some embodiments, the amplified fragments comprise at least one 20 base pair sequence comprising at least one, at least two, at least three, or more than three CpG dinucleotides. Amplification can be performed using a set of primer oligonucleotides according to the present disclosure, and can use a thermostable polymerase. Amplification of several DNA segments can be performed simultaneously in the same reaction vessel. In some embodiments, two or more fragments are amplified simultaneously. Amplification can be performed using, for example, polymerase chain reaction (PCR).

[0247] Primers designed to target these sequences can exhibit some degree of bias towards converted methylated sequences. In some embodiments, PCR primers are designed to be methylation specific for targeting methylation sequencing applications. This can allow for higher sensitivity in some applications. For example, primers can be designed to include a discriminant nucleotide (specific to methylated sequences after bisulfite conversion) positioned to achieve optimal discrimination (e.g., in PCR applications). The discriminant can be at the 3' end or penultimate position.

[0248] In some embodiments, primers are designed to amplify DNA fragments of 75 to 350 bp in length. This is a general size range known for circulating DNA, and optimizing primer design to account for target size can improve sensitivity of the method according to the present example. Primers can be designed to amplify regions of about 50 to 200, about 75 to 150, or about 100 or 125 bp in length.

[0249] In some embodiments of the methods described herein, methylation-specific primer oligonucleotides can be used to detect the methylation status of preselected CpG positions in a nucleic acid sequence by amplification-based methods. Amplification of bisulfite-treated DNA using methylation status-specific primers allows for the discrimination of methylated versus unmethylated nucleic acids. MSP primer pairs contain at least one primer that hybridizes to a converted CpG dinucleotide. Thus, the sequence of the primer includes at least one CpG, TpG, or CpA dinucleotide. MSP primers specific for non-methylated DNA contain a “T” at the 3’ position of the C in the CpG. Thus, the base sequence of the primer can require inclusion of a sequence of at least 18 nucleotides in length that hybridizes to the pre-processed nucleic acid sequence and its complement, wherein the base sequence of the oligomer includes at least one CpG, TpG, or CpA dinucleotide. In some embodiments, MSP primers include 2 to 5 CpG, TpG, or CpA dinucleotides. In some embodiments, the dinucleotides are located within the 3’ half of the primer, e.g., for a primer of 18 bases in length, the specified dinucleotides are located within the first 9 bases from the 3’ end of the molecule. In addition to the CpG, TpG, or CpA dinucleotides, the primers can also include several methylated converted bases (e.g., cytosine to thymine, or guanine to adenine on the hybridized strand). In some embodiments, the primers are designed to include no more than 2 cytosine or guanine bases.

[0250] In some embodiments, each region is segmented for amplification with multiple primer pairs. In some embodiments, the segments do not overlap. The segments can be directly adjacent or spaced apart (e.g., by up to 10, 20, 30, 40, or 50 bp). As target regions (including CpG islands, CpG shores, and / or CpG shelves) are typically longer than 75 to 150 bp, the present example allows for the assessment of the methylation status of more (or all) sites spanning a given target region.

[0251] Primers can be designed for target regions using suitable tools such as Primer3, Primer3Plus, Primer-BLAST, and the like. As discussed, bisulfite conversion results in the conversion of cytosine to uracil, and 5’-methylcytosine to thymine. Thus, primer positioning or targeting can take advantage of the methylation sequence following bisulfite conversion, depending on the degree of methylation specificity desired.

[0252] The amplified target region is designed to have at least 10 CpG dinucleotide methylation sites. However, in some examples, it can be advantageous to amplify a region having more than 10 CpG methylation sites. For example, a 300 bp long sequence read can have about 10, 20, 30, 40, or 50 CpG methylation sites that are methylated in a nucleic acid sample associated with a colon cell proliferative disorder. In various examples, the methylation regions identified in Tables 1-11 can have at least 25, 50, 100, 200, 300, 400, or 500 CpG methylation sites that are methylated in a nucleic acid sample associated with a colon cell proliferative disorder. In some embodiments, primers are designed to amplify a DNA fragment comprising 3 to 20 CpG methylation sites in the targeted region. Overall, this approach allows for more methylation sites to be interrogated in a single sequencing read and provides additional certainty (to exclude false positives) because multiple consistent methylation can be detected in a single sequencing read. In some embodiments, the tumor signal comprises more than two methylation regions selected from Tables 1-11. In this example, detecting multiple tumor signals can increase the confidence of tumor detection. Such signals can be at the same site or at different sites. In some embodiments, detection of more than one tumor signal at the same region is indicative of a tumor.

[0253] In some embodiments, the number of CpG sites in the identified methylation regions can be modeled between two populations with different characteristics of colon cell proliferative disorders to identify a methylation threshold, where the number of CpG sites in a region exceeding the threshold is indicative of a colon cell proliferative disorder.

[0254] In various examples, the number of CpG sites in the identified methylation regions that are indicative of colorectal cancer is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18, where the presence of methylated CpGs if exceeding this identified number is indicative of colorectal cancer and can be used as an input feature for a machine learning model that is used as a classifier to stratify a population into healthy individuals and colorectal cancer individuals.

[0255] In the present example, detection of multiplexed tumor signals indicative of methylation at the same site in the genome can increase the confidence of tumor detection. Detection of methylation at adjacent sites in the genome, even if the signals are from different sequencing reads, can increase the confidence of tumor detection. This reflects another type of signal consistency. In some embodiments, detection of tumor signals that are adjacent or overlapping in at least two different sequencing reads is indicative of a tumor. In some embodiments, the adjacent or overlapping tumor signals are within the same CpG island. In some embodiments, detection of 3 to 34 proximal methylation sites in a cell-free DNA fragment is indicative of a tumor. In some embodiments, detection of 3 to 34 methylation CpG sites in a fragment is used to identify a threshold to distinguish a population of individuals with a certain characteristic (e.g., healthy, disease, or stage of disease). In some embodiments, detection of about 4 to 10, about 4 to 15, about 10 to 20, about 15 to 20, about 15 to 25, about 20 to 25, about 20 to 34, about 25 to 34, or about 30 to 34 methylation proximal CpG sites in a read fragment is used to determine a threshold to distinguish a population of individuals with a certain characteristic (e.g., healthy, disease, or stage of disease). As used herein, the term "proximal CpG site" refers to a CpG site in a cell-free nucleic acid sample that is adjacent to or between 2 to 10 CpG sites on the same nucleic acid fragment.

[0256] In some embodiments, amplification is performed using more than 100 primer pairs. Amplification can be performed using about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, or more primer pairs. In some embodiments, amplification is multiplexed amplification. Multiplexed amplification allows collection of large amounts of methylation information from many target regions of the genome in parallel, even from cfDNA samples where DNA is typically not abundant. Multiplexing can be scaled to a platform, such as the Ion Torrent® PGM platform. In some embodiments, amplification is nested amplification. Nested amplification can improve sensitivity and specificity.

[0257] In addition, another fast and robust approach for parallel examination of multiple methylation sequences is known as synchronous targeted methylation sequencing (sTM-Seq). Key features of this technology include the elimination of the need for large amounts of high molecular weight DNA, and nucleotide-specific discrimination of 5-methylcytosine (5mC) from 5-hydroxymethylcytosine (5hmC). In addition, sTM-Seq is scalable and can be used to investigate multiple loci in dozens of samples in a single sequencing run. Web-based software and universal primers for multipurpose barcoding, library preparation, and custom sequencing are freely available, making sTM-Seq affordable, efficient, and broadly applicable (e.g., as described by Asmus, N. et al., Curr Protoc Hum Genet. 2019 Apr; 101(1), the contents of which are incorporated herein by reference).

[0258] Generally, the methods and systems provided herein are useful for preparing cell-free polynucleotide sequences for downstream application sequencing reactions. In some embodiments, the sequencing method is classical Sanger sequencing. Sequencing methods can include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq Digital Gene Expression Next-generation sequencing, single molecule sequencing by synthesis (SMSS) Massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.

[0259] Pyrosequencing can refer to a real-time sequencing technique based on photometric detection of pyrophosphate release following nucleotide incorporation, suitable for simultaneous analysis and quantification of methylation degree at several CpG positions. After genomic DNA transformation, the region of interest is amplified with polymerase chain reaction (PCR), in which one of the two primers is biotinylated. The template generated by PCR is presented as single-stranded, and pyrosequencing primers are annealed for quantitative analysis of CpG positions. After bisulfite treatment and PCR, each methylation degree at each CpG position in the sequence is determined by the ratio of T to C signals, reflecting the proportion of unmethylated to methylated cytosine at each CpG site in the original sequence.

[0260] V. Classifiers, machine learning models, and systems

[0261] In various examples, methylation sequencing features are used as input datasets for trained algorithms (e.g., machine learning models or classifiers) to find correlations between sequence composition and patient groupings. Examples of such patient groupings include presence, stage, subtype, responder vs. non-responder, and progressor vs. non-progressor of a disease or condition. In various examples, a feature matrix is generated to compare samples obtained from individuals with known conditions or characteristics. In some embodiments, samples are obtained from healthy individuals or individuals without any known indication and samples are obtained from patients known to have cancer.

[0262] As used herein, with respect to machine learning and pattern recognition, the term "feature" generally refers to a single measurable characteristic or trait of an observed phenomenon. The concept of "feature" is related to the concept of explanatory variable used in statistical techniques such as, but not limited to, linear regression and logistic regression. Features are typically numeric, but structural features such as strings and graphs are used in syntactic pattern recognition.

[0263] As used herein, the term "input feature" (or "feature") generally refers to a variable, such as a condition, sequence content (e.g., mutation), proposed data collection operation, or proposed treatment, used by a trained algorithm (e.g., model or classifier) to predict an output classification (label) for a sample. The value of the variable can be determined for one sample and used to determine the classification.

[0264] In various examples, input features of genetic data include alignment variables related to alignment of sequence data (e.g., sequence reads) to a genome, and non-alignment variables, e.g., related to sequence content of sequence reads, measurements of proteins or autoantibodies, or average methylation levels of genomic regions. Input features can be gene features, such as V-plot metrics, FREE-C deconvolution, chromatin accessibility, and cfDNA measurements at transcription start sites. Indicators that can be used in methylation analysis include, but are not limited to, base-by-base methylation percentages for CpG, CHG, CHH, conversion efficiency (100 - average methylation percentage for CHH), hypomethylated segments, methylation levels (overall average methylation for CPG, CHH, CHG, segment length, segment midpoint, and methylation levels in one or more genomic regions such as chrM, LINE1, or ALU), number of methylated CpGs per segment, fraction of CpG methylation per segment, fraction of CpG methylation per region, fraction of CpG methylation in panel, dinucleotide coverage (normalized dinucleotide coverage), coverage uniformity (unique CpG sites at 1x and 10x average genome coverage (for S4 runs)), overall average CpG coverage (depth), and average coverage at CpG islands, CGI shores, and CGI shores. These indicators can be used as feature inputs for machine learning methods and models.

[0265] For multiple assays, the system identifies a set of features to input into a trained algorithm (e.g., a machine learning model or classifier). The system analyzes each molecular class and forms a feature vector from the measurements. The system inputs the feature vector into the machine learning model and obtains an output classification of whether the biological sample has a specified characteristic.

[0266] In some embodiments, the machine learning model outputs a classifier that is able to distinguish between two or more groups or classes of individuals or features of a population of individuals. In some embodiments, the classifier is a trained machine learning classifier.

[0267] In some embodiments, information loci or features of biomarkers in tumor tissue are analyzed to form a profile. A receiver operator characteristic (ROC) curve can be generated by plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any additional biomedical information item) in distinguishing between two populations (e.g., individuals who respond to a therapeutic agent and those who do not). In some embodiments, the feature data across the entire population (e.g., cases and controls) is sorted in ascending order based on individual feature values.

[0268] In various examples, the specified characteristic is selected from the group consisting of health and cancer, disease subtype, disease stage, progressor and non-progressor, and responder and non-responder.

[0269] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of adenoma (adenomatous polyps), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal carcinoma, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal carcinoma.

[0270] A. Data Analysis

[0271] In some examples, the present disclosure provides a system, method, or kit in which data analysis can be implemented in a software application, computing hardware, or both. In various examples, the analysis application or system comprises at least one data receiving module, one data pre-processing module, one data analysis module (which can operate on one or more types of genomic data), one data interpretation module, or one data visualization module. In some embodiments, the data receiving module can comprise a computer system that connects laboratory hardware or instruments with a computer system that processes laboratory data. In some embodiments, the data pre-processing module can comprise a hardware system or computer software that performs operations on data in preparation for analysis. Examples of operations that can be applied to data in the pre-processing module include affine transformations, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module can be specialized to analyze genomic data from one or more genomic materials, for example, assembled genomic sequences can be taken and probabilistic and statistical analyses performed to identify abnormal patterns that are associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analytical methods, for example, drawn from statistics, mathematics, or biology, to support understanding of the relationship between the identified abnormal patterns and a health condition, functional state, prognosis, or risk. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to create visualizations of the data that can facilitate understanding or interpretation of the results.

[0272] In various examples, a machine learning method is applied to distinguish between samples in a population of samples. In some embodiments, a machine learning method is applied to distinguish between healthy and advanced disease (e.g., adenoma) samples.

[0273] In some embodiments, the one or more machine learning operations used to train the prediction engine include one or more of: a generalized linear model, a generalized additive model, a non-parametric regression operation, a random forest classifier, a spatial regression operation, a Bayesian regression model, a time series analysis, a Bayesian network, a Gaussian network, a decision tree learning operation, an artificial neural network, a recurrent neural network, a convolutional neural network, a reinforcement learning operation, a linear or non-linear regression operation, a support vector machine, a clustering operation, and a genetic algorithm operation.

[0274] In various examples, the computer processing method is selected from logistic regression, multivariate linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neural fuzzy, and artificial neural networks.

[0275] In some examples, the methods disclosed herein can include computational analysis of nucleic acid sequencing data from a sample of an individual or a plurality of individuals.

[0276] B. Classifier Generation

[0277] In one aspect, the disclosed systems and methods provide a classifier that is generated based on feature information derived from methylation sequence analysis of cfDNA biological samples. The classifier forms part of a prediction engine for distinguishing groups in a population based on sequence features identified in biological samples, such as cfDNA.

[0278] In some embodiments, the classifier is created by normalizing sequence information by formatting similar portions of the sequence information into a uniform format and a uniform scale; storing the normalized sequence information in a columnar database; training a prediction engine by applying one or more machine learning operations to the stored normalized sequence information to predict a combination of one or more features for a particular population; applying the prediction engine to field information accessed to identify individuals associated with a grouping; and dividing the individuals into the grouping.

[0279] In some embodiments, the hierarchy is created by normalizing the sequence information by formatting similar portions of the sequence information into a uniform format and a uniform scale; storing the normalized sequence information in a columnar database; training a prediction engine by applying one or more machine learning operations to the stored normalized sequence information to predict a combination of one or more features for a particular group; applying the prediction engine to the accessed field information to identify individuals related to the grouping; and dividing the individuals into the grouping.

[0280] Specificity, as used herein, generally refers to the "probability of a negative test result in an individual without the disease." It can be calculated as the number of disease-free individuals with a negative test result divided by the total number of disease-free individuals.

[0281] In various examples, the model, classifier, or predictive test has a specificity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0282] Sensitivity, as used herein, generally refers to the "probability of a positive test result in an individual with the disease." It can be calculated as the number of diseased individuals with a positive test result divided by the total number of diseased individuals.

[0283] In various examples, the model, classifier, or predictive test has a sensitivity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0284] Positive predictive value, as used herein, generally refers to the "probability of a positive test result being correct." It can be calculated as the number of true positive test results divided by the total number of positive test results.

[0285] In various examples, the model, classifier, or predictive test has a positive predictive value of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0286] Negative predictive value, as used herein, generally refers to the "probability of a negative test result being correct." It can be calculated as the number of true negative test results divided by the total number of negative test results.

[0287] In various examples, the model, classifier, or prediction test has a negative predictive value of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0288] C. Digital processing device

[0289] In some examples, the subject matter described herein can include a digital processing device or uses thereof. In some examples, the digital processing device can include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform the functions of the device. In some examples, the digital processing device can include an operating system configured to execute executable instructions.

[0290] In some examples, the digital processing device is optionally connected to a computer network. In some examples, the digital processing device is optionally connected to the Internet. In some examples, the digital processing device is optionally connected to a cloud computing facility. In some examples, the digital processing device is optionally connected to an intranet. In some examples, the digital processing device is optionally connected to a data storage device.

[0291] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, notebook computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, handheld computers, Internet appliances, mobile smartphones, and tablet computers. Suitable tablet computers can include, for example, those having a booklet, notepad, and convertible configurations.

[0292] In some examples, the digital processing device can include an operating system configured to execute executable instructions. For example, the operating system can include software, including programs and data, that manages the device’s hardware and provides services for the execution of applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, Linux, Mac OS X Windows and Non-limiting examples of suitable personal computer operating systems include Mac OS and UNIX-like operating systems, such as In some examples, the operating system can be provided by a cloud computing, and the cloud computing resources can be provided by one or more service providers.

[0293] In some examples, the device can include a storage and / or memory device. The storage and / or memory device can be one or more physical devices used to store data or programs on a temporary or permanent basis. In some examples, the device can be a volatile memory and require power to maintain stored information. In some examples, the device is a non-volatile memory and retains stored information when the digital processing device is not powered. In some examples, the non-volatile memory can include flash memory. In some examples, the non-volatile memory can include dynamic random access memory (DRAM). In some examples, the non-volatile memory can include ferroelectric random access memory (FRAM). In some examples, the non-volatile memory can include phase-change random access memory (PRAM).

[0294] In some examples, the device can be a storage device including, for example, a CD-ROM, a DVD, a flash memory device, a disk drive, a tape drive, an optical disk drive, and cloud computing-based storage. In some examples, the storage and / or memory device can be a combination of devices such as those disclosed herein. In some examples, the digital processing device can include a display to send visual information to a user. In some examples, the display can be a cathode ray tube (CRT). In some examples, the display can be a liquid crystal display (LCD). In some examples, the display can be a thin-film transistor liquid crystal display (TFT-LCD). In some examples, the display can be an organic light-emitting diode (OLED) display. In some examples, the OLED display can be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some examples, the display can be a plasma display. In some examples, the display can be a video projector. In some examples, the display can be a combination of devices such as those disclosed herein.

[0295] In some examples, the digital processing device can include an input device to receive information from a user. In some examples, the input device can be a keyboard. In some examples, the input device can be a pointing device including, for example, a mouse, a trackball, a trackpad, a joystick, a game controller, or a stylus. In some examples, the input device can be a touch screen or a multi-touch screen. In some examples, the input device can be a microphone to capture voice or other sound input. In some examples, the input device can be a video camera to capture motion or visual input. In some examples, the input device can be a combination of devices such as those disclosed herein.

[0296] D. Non-transitory computer-readable storage medium

[0297] In some examples, the subject matter disclosed herein can include one or more non-transitory computer-readable storage media having program code stored on the media, the program code comprising instructions executable by an optional network digital processing apparatus. In some examples, the computer-readable storage media can be a tangible component of the digital processing apparatus. In some examples, the computer-readable storage media is optionally removable from the digital processing apparatus. In some examples, the computer-readable storage media can include, for example, a CD-ROM, a DVD, a flash memory device, a solid-state memory, a disk drive, a tape drive, an optical disk drive, a cloud computing system and service, and the like. In some examples, the program and instructions can be encoded on the media in a non-transitory, semi-transitory, or non-transient manner.

[0298] E. Computer System

[0299] The present disclosure provides a computer system programmed to implement the methods described herein. Figure 1 A computer system 101 is shown that is programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, and reference sequences. The computer system 101 can process various aspects of the patient data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system 101 can be an electronic device of a user or a computer system located remotely from the electronic device. The electronic device can be a mobile electronic device.

[0300] The computer system 101 includes a central processing unit (CPU, also "processor" and "computer processor" herein) 105, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 101 also includes memory or memory location 110 (e.g., random access memory, read only memory, flash memory), electronic storage unit 115 (e.g., hard disk), communication interface 120 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 125, such as cache, other memory, data storage, and / or electronic display adapters. The memory 110, storage unit 115, interface 120, and peripheral devices 125 are in communication with the CPU 105 through a communication bus (solid lines), such as a motherboard. The storage unit 115 can be a data storage unit (or data repository) for storing data. The computer system 101 can be operatively coupled to a computer network ("network") 130 by means of the communication interface 120. The network 130 can be the Internet, an internet and / or an extranet, or an intranet and / or extranet that in turn can use a private communication protocol to communicate with an Internet, an internet, or the like. The network 130 can include one or more computer servers, which can implement a distributed computing solution. In some instances, the network 130 can implement a point-to-point network, which can enable devices coupled to the computer system 101 to behave as a client-server.

[0301] The CPU 105 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 110. The instructions can be directed to the CPU 105, which can subsequently program or otherwise configure the CPU 105 to implement methods of the present disclosure. Examples of operations performed by the CPU 105 can include fetch, decode, execute, and writeback.

[0302] The CPU 105 can be part of a circuit, such as an integrated circuit. One or more other components of the system 101 can be included in the circuit. In some instances, the circuit is an application specific integrated circuit (ASIC).

[0303] The storage unit 115 can store files, such as drivers, libraries and saved programs. The storage unit 115 can store user data, e.g., user preferences and user programs. The computer system 101 in some instances can include one or more additional data storage units that are external to the computer system 101, such as located on a remote server that is in communication with the computer system 101 through an intranet or the Internet.

[0304] Computer system 101 can communicate with one or more remote computer systems through network 130. For instance, computer system 101 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., desktop PC), slate / tablet PCs (e.g., Apple® iPad, Apple® iPad®, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. The user can access the computer system 101 via network 130. The methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 101, such as, for example, on the memory 110 or electronic storage 115. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 105. In some examples, the code can be retrieved from the storage 115 and stored on the memory 110 for ready access by the processor 105. In some examples, the electronic storage 115 can be precluded, and machine-executable instructions are stored on memory 110. The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled for use at runtime. The code can be provided in a programming language that can be selected to enable the code to execute on one or more specific processors, such as the processor 105.

[0305] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled for use at runtime. The code can be provided in a programming language that can be selected to enable the code to execute on one or more specific processors, such as the processor 105.

[0306]

[0307] ​​Aspects of the systems and methods provided herein, such as computer system 101, can be embodied in programming. Various aspects of the technology can be thought of as "products" or "articles of manufacture" typically in the form of machine (or processor) executable code and / or associated data that is carried or otherwise transported by a type of machine readable medium. Machine-executable code can be stored on or in one or more of such tangible, machine-readable media as aforementioned "storage" types, which are for the non-transitory storage of information. It has been

[0308] Accordingly, a machine readable medium, such as a computer-executable code, can take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optic cables, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards, paper tape, any other physical storage medium that can be used to store or transmit information, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transmitted or otherwise propagating through a cable or link, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0309] The computer system 101 can include or be in communication with an electronic display 135, which can comprise a user interface (UI) 140 for providing, for example, nucleic acid sequences, concentrated nucleic acid samples, methylation profiles, expression profiles, and analyses of methylation or expression profiles. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces.

[0310] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 105. For example, the algorithms can store, process, identify, or interpret patient data, biological data, biological sequences, and reference sequences.

[0311] While certain examples of methods and systems have been shown and described, it will be understood by those skilled in the art that these are presented by way of example only and are not intended to limit the scope of the disclosure herein. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the scope of the disclosure herein. Moreover, it should be understood that all aspects of the methods and systems described herein are not limited to the specific descriptions, configurations, or relative proportions set forth herein, which depend on a variety of conditions and variables. The descriptions are intended to include all such alternatives, modifications, variations, or equivalents.

[0312] In some examples, the subject matter disclosed herein can include at least one computer program or uses thereof. The computer program can be a sequence of instructions written in a programming language that can be executed by a CPU, GPU, or TPU of a digital processing device. The computer program can be loaded into the digital processing device from a tangible computer readable storage medium, which can be, for example, a computer diskette, a CD ROM, a solid state memory device, or an optical disk such as a DVD or a Blu-ray disk. The computer program can be stored on the digital processing device, loaded into the digital processing device, and / or executed by the digital processing device. The computer program can be executed by the digital processing device in a manner that causes the digital processing device to perform a process described herein.

[0313] In various environments, the functions of the computer readable instructions can be combined or distributed as desired. In some examples, the computer program can include one sequence of instructions. In some examples, the computer program can include multiple sequences of instructions. In some examples, the computer program can be provided by one location. In some examples, the computer program can be provided by multiple locations. In some examples, the computer program can include one or more software modules. In some examples, the computer program can include, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or a combination thereof.

[0314] In some examples, the computer processing method is a supervised machine learning method including, for example, regression, support vector machines, tree-based methods, and networks.

[0315] In some examples, the computer processing method is an unsupervised machine learning method including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0316] In some examples, the computer processing method is an unsupervised machine learning method including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0317] F. Database

[0318] In some examples, the subject matter disclosed herein can include one or more databases, or uses of the databases, to store patient data, biological data, biological sequences, or reference sequences. Reference sequences can be derived from the databases. In view of the disclosure provided herein, many databases can be suitable for storing and retrieving sequence information. In some examples, suitable databases can include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some examples, the databases can be internet-based. In some examples, the databases can be web-based. In some examples, the databases can be cloud-computing based. In some examples, the databases can be based on one or more local computer storage devices.

[0319] In an aspect, the disclosure provides a non-transitory computer readable medium comprising instructions to instruct a processor to perform the methods described herein.

[0320] In an aspect, the disclosure provides a computing device comprising a computer readable medium.

[0321] In another aspect, the disclosure provides a system for classifying a biological sample, comprising: a) a receiver to receive a plurality of training samples, each of the plurality of training samples having a plurality of classes of molecules, wherein each of the plurality of training samples comprises one or more known labels, b) a feature module to identify a set of operable features corresponding to assays to input into a machine learning model for each of the plurality of training samples, wherein the set of features correspond to properties of the molecules in the plurality of training samples, wherein for each of the plurality of training samples, the system is operable to cause the plurality of classes of molecules in the training sample to undergo a plurality of different assays to obtain a set of measurements, wherein each set of measurements is from one assay of one class of molecules in the training sample, wherein the plurality of sets of measurements are obtained for the plurality of training samples, c) an analysis module to analyze the sets of measurements to obtain a training vector for the training sample, wherein the training vector comprises feature values for N sets of features corresponding to the assays, each feature value corresponding to a feature and comprising one or more measurements, wherein the training vector is formed using at least one feature from at least two of the N sets of features corresponding to a first subset of the plurality of different assays, d) a label module to inform the system about the training vector using parameters of the machine learning model to obtain an output label for the plurality of training samples, e) a comparator module to compare the output label to the known label of the training sample, f) a training module to iteratively search for an optimal value of the parameters as part of training the machine learning model based on the comparison of the output label to the known label of the training sample, and g) an output module to provide the parameters of the machine learning model and the set of features of the machine learning model.

[0322] VI. Methods of classifying subjects in a population

[0323] The disclosed methods aim to determine genetic and / or epigenetic parameters of genomic DNA associated with a colon cell proliferative disorder through analysis of cfDNA in a subject. The methods are used to improve diagnosis, treatment, and monitoring of colon cell proliferative disorders, more specifically, by improving identification and differentiation between stages or sub-classes of the disorder and genetic predisposition to the disorder.

[0324] In some embodiments, the methods comprise analyzing methylation status of CpG islands, CpG shores, or CpG shelves.

[0325] In some embodiments, the methods comprise analyzing methylation status, hemimethylation status, hypermethylation status, or hypomethylation status of cell-free nucleic acids in the biological sample.

[0326] In one aspect, the disclosure provides a method for detecting a colon cell proliferative disorder that can be applied to cell-free samples, e.g., to detect cell-free circulating colon cell proliferative disorder DNA. The method utilizes detection of methylation signal in a single sequencing read as a basic “positive” colon cell proliferative disorder signal.

[0327] In some embodiments, the colon cell proliferative disorder is selected from the group consisting of: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal carcinoma, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, the colon cell proliferative disorder comprises colorectal carcinoma.

[0328] In one aspect, the disclosure provides a method for detecting a colon cell proliferative disorder comprising: extracting DNA from a cell-free sample obtained from a subject, transforming at least a portion of the DNA for methylation sequencing, amplifying methylation regions resulting from the transformed DNA in cancer, generating sequencing reads from the amplified regions, and detecting a signal of a colon cell proliferative disorder comprising at least one, at least two, at least three, or more than three methylation regions within a cancer panel to obtain an input feature, which is input into a machine learning model to obtain a classifier that is able to distinguish between two groups of subjects (e.g., healthy vs. cancer, disease stage, advanced adenoma vs. cancer).

[0329] The trained machine learning methods, models, and discriminative classifiers described herein can be applied to various medical applications, including cancer detection, diagnosis, and treatment responsiveness. As the models can be trained with individual metadata and analyte-derived features, the applications can be customized to stratify individuals in a population and direct treatment decisions accordingly.

[0330] Diagnosis

[0331] The methods and systems provided herein can perform predictive analysis using artificial intelligence-based methods to analyze data obtained from a subject (patient) to generate a diagnostic output for a subject with cancer (e.g., colorectal cancer). For example, the application can apply a predictive algorithm to the acquired data to generate a diagnosis for a subject with cancer. The predictive algorithm can include an artificial intelligence-based predictor, such as a machine learning-based predictor, configured to process the acquired data to generate a diagnosis for a subject with cancer.

[0332] The machine learning predictor can be trained using a dataset, e.g., a dataset generated from methylation assays performed on individual biological samples from one or more cohorts of subjects with cancer using the signature panels described herein as input, and known diagnostic (e.g., stage and / or tumor fraction) outcomes for the subjects as output for the machine learning predictor.

[0333] The training dataset (e.g., a dataset generated from methylation assays performed on individual biological samples using the signature panels described herein) can be generated from, e.g., one or more cohorts of subjects having common characteristics (features) and outcomes (labels). The training dataset can include a set of features and labels corresponding to diagnostic-related features. The features can include some characteristic, e.g., like certain ranges or categories measured by cfDNA assays, such as counts of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap or fall in a set of bins (genomic windows) of a reference genome. For example, a set of features collected from a given subject at a given point in time can collectively serve as a diagnostic signature that can be indicative of the subject having an identified cancer at the given point in time. The characteristics can also include labels that are indicative of a subject’s diagnostic outcome, such as one or more cancers.

[0334] The labels can include outcomes, e.g., known diagnostic (e.g., stage and / or tumor fraction) outcomes for the subjects. The outcomes can include characteristics related to a subject’s cancer. For example, the characteristics can be indicative of a subject having one or more cancers.

[0335] A training set (e.g., training dataset) can be selected by random sampling of one dataset corresponding to one or more subject sets (e.g., retrospective and / or prospective patient cohorts with or without one or more cancers). Alternatively, a training set (e.g., training dataset) can be selected by proportional sampling of one dataset corresponding to one or more subject sets (e.g., retrospective and / or prospective patient cohorts with or without one or more cancers). A training set can be balanced among datasets corresponding to one or more subject sets (e.g., patients from different clinical sites or trials). A machine learning predictor can be trained until certain predetermined accuracy or performance conditions are met, such as having a minimum expected value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure can correspond to a prediction of a diagnosis, stage, or tumor fraction of one or more cancers of a subject.

[0336] Examples of diagnostic accuracy measures can include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC) of a receiver operating characteristic (ROC) curve corresponding to diagnostic accuracy of detecting or predicting a cancer (e.g., colorectal cancer).

[0337] In one aspect, the present disclosure provides a method of using a classifier capable of distinguishing between a population of individuals, comprising: a) performing an analysis of a plurality of classes of molecules in a biological sample, wherein the analysis provides a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to a characteristic of each of the plurality of classes of molecules input into a machine learning or statistical model; c) preparing a feature vector of feature values from each of the plurality of sets of measurements, each feature value corresponding to one feature of the set of features and comprising one or more measurements, wherein the feature vector comprises at least one feature value obtained using each of the plurality of sets of measurements; d) loading into a memory of a computer system: a trained machine learning model comprising a classifier, the trained machine learning model trained using training vectors obtained from training biological samples, a first subset of training biological samples identified as having a specified characteristic, and a second subset of training biological samples identified as not having the specified characteristic; and e) applying the trained machine learning model to the feature vector to obtain an output classification of whether the biological sample has the specified characteristic, thereby distinguishing between the population of individuals having the specified characteristic.

[0338] In one aspect, the present disclosure provides a method for using a hierarchy capable of distinguishing between a population of individuals, comprising: a) performing an analysis on a plurality of classes of molecules in a biological sample, wherein the analysis provides a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to a characteristic of each of the plurality of classes of molecules inputted into a machine learning or statistical model; c) preparing a feature vector of feature values from each of the plurality of sets of measurements, each feature value corresponding to one feature of the set of features and comprising one or more measurements, wherein the feature vector comprises at least one feature value obtained using each of the plurality of sets of measurements; d) loading into a memory of a computer system: a trained machine learning model comprising a classifier, the trained machine learning model trained using training vectors obtained from training biological samples, a first subset of training biological samples identified as having a specified characteristic, and a second subset of training biological samples identified as not having the specified characteristic; and e) applying the trained machine learning model to the feature vector to obtain an output classification of whether the biological sample has the specified characteristic, thereby distinguishing between a population of individuals having the specified characteristic.

[0339] In one aspect, the present disclosure provides a method for using a hierarchy capable of distinguishing between a population of individuals, comprising: a) detecting methylation signals in one or more first patient samples in single sequencing reads of a pre-selected genomic region, b) the methylation signals affecting a hierarchy of data output, thereby affecting a machine learning model, and c) a second patient sample using the affected hierarchy to detect methylation signals.

[0340] In some embodiments, the pre-selected genomic region is selected from two or more methylation genomic regions in Tables 1-11, three or more methylation genomic regions in Tables 1-11, four or more methylation genomic regions in Tables 1-11, five or more methylation genomic regions in Tables 1-11, six or more methylation genomic regions in Tables 1-11, seven or more methylation genomic regions in Tables 1-11, eight or more methylation genomic regions in Tables 1-11, nine or more methylation genomic regions in Tables 1-11, ten or more methylation genomic regions in Tables 1-11, eleven or more methylation genomic regions in Tables 1-11, twelve or more methylation genomic regions in Tables 1-11, or thirteen or more methylation genomic regions in Tables 1-11.

[0341] In another aspect, the present disclosure provides a method for identifying a cancer in a subject, comprising: a) providing a biological sample containing cell-free nucleic acid (cfNA) molecules from the subject; b) subjecting the cfNA molecules from the subject to methylation conversion and sequencing to generate a plurality of cfNA sequencing reads; c) aligning the plurality of cfNA sequencing reads to a reference genome; d) generating a quantitative measure of the plurality of cfNA sequencing reads on each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 different regions, each of the at least about 10 different regions comprising at least a portion of a gene selected from the methylation regions in the signature panel described herein; and e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has the cancer.

[0342] In some examples, the at least about 10 different regions comprise at least about 20 different regions, each of the at least about 20 different regions comprising at least a portion of a methylation region identified in Tables 1-11. In some examples, the at least about 10 different regions comprise at least about 30 different regions, each of the at least about 30 different regions comprising at least a portion of a methylation region identified in Tables 1-11.

[0343] As another example, such a predetermined condition can be a specificity for predicting a colon cell proliferative disorder comprising a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0344] As another example, such a predetermined condition can be a positive predictive value (PPV) for predicting a colon cell proliferative disorder comprising a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0345] As another example, such predetermined condition can be a negative predictive value (NPV) for predicting a colon cell proliferative disorder comprising a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0346] As another example, such predetermined condition can be an area under the curve (AUC) of a receiver operating characteristic (ROC) curve for predicting a colon cell proliferative disorder comprising a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0347] Therapeutic responsiveness

[0348] The predictive classifiers, systems, and methods described herein can be used to classify a population of individuals for a variety of clinical applications (e.g., based on performing methylation assays on individual biological samples using the signature panels described herein). Examples of such clinical applications include detecting early stage cancer, diagnosing cancer, classifying cancer into a particular disease stage, determining responsiveness or resistance to a therapeutic agent for treating cancer.

[0349] The methods and systems described herein can be applied to the characterization of colon cell proliferative disorders, such as staging and classification. Thus, the combination of analytes and assays can be used in the present systems and methods to predict responsiveness of different cancer types in different tissues to cancer therapeutic agents, and to classify individuals according to therapeutic responsiveness. In some embodiments, the classifiers described herein are capable of stratifying a group of individuals into treatment responders and non-responders.

[0350] The present disclosure also provides a method for determining drug targets (e.g., genes relevant or important to a particular class) for a target condition or disease, comprising: assessing gene expression levels of at least one gene in a sample obtained from an individual; and using a neighborhood analysis program, determining genes relevant to a classification of the sample, thereby determining one or more drug targets relevant to the classification.

[0351] The present disclosure also provides a method for determining the efficacy of a drug designed to treat a class of diseases, comprising: obtaining a sample from an individual having the class of diseases; subjecting the sample to the drug; assessing the gene expression level of at least one gene in the drug-exposed sample; and using a computer model established with a weighted voting scheme, classifying the drug-exposed sample into a class of diseases as a function of the relative gene expression level of the sample to the model.

[0352] The present disclosure also provides a method for determining the efficacy of a drug designed to treat a class of diseases, wherein the individual has been subjected to the drug, the method comprising obtaining a sample from the individual subjected to the drug; assessing the gene expression level of at least one gene in the sample; and using a model established with a weighted voting scheme, classifying the sample into a class of diseases comprising assessing the gene expression level of the sample compared to the gene expression level of the model.

[0353] The present disclosure also provides a method for determining whether an individual belongs to a class of phenotypes (e.g., intelligence, response to treatment, length of life, likelihood of viral infection, or obesity), comprising: obtaining a sample from the individual; assessing the gene expression level of at least one gene in the sample; and using a model established with a weighted voting scheme, classifying the sample into a class of diseases comprising assessing the gene expression level of the sample compared to the gene expression level of the model.

[0354] In one aspect, the systems and methods described herein relating to classifying a population based on treatment responsiveness refer to cancer treated with chemotherapy agents in the classes of, but not limited to, DNA damaging agents, DNA repair targeting therapies, DNA damage signaling inhibitors, DNA damage induction cell cycle arrest inhibitors, and inhibitors of processes that indirectly cause DNA damage. Each of these chemotherapy agents can be considered a "DNA damaging therapeutic agent" as the term is used herein.

[0355] Based on the patient's analyte data, the patient can be classified into high risk and low risk patient groups, such as patients with high or low risk of clinical recurrence, and the results can be used to determine the course of treatment. For example, a patient determined to be a high risk patient can receive adjuvant chemotherapy after surgery. For a patient deemed to be a low risk patient, adjuvant chemotherapy can be discontinued after surgery. Thus, the present disclosure provides, in certain aspects, a method of preparing a colon cancer tumor gene expression profile indicative of risk of recurrence.

[0356] In various examples, the classifiers described herein are capable of stratifying a population of individuals between responders and non-responders to a treatment.

[0357] In another aspect, the methods disclosed herein can be applied to clinical applications involving cancer detection or monitoring.

[0358] In some embodiments, the methods disclosed herein can be used to determine and / or predict response to therapy.

[0359] In some embodiments, the methods disclosed herein can be used to monitor and / or predict tumor burden.

[0360] In some embodiments, the methods disclosed herein can be used to detect and / or predict post-surgical residual tumor.

[0361] In some embodiments, the methods disclosed herein can be used to detect and / or predict minimal residual disease after therapy.

[0362] In some embodiments, the methods disclosed herein can be used to detect and / or predict recurrence.

[0363] In one aspect, the methods disclosed herein can be used as a secondary screening.

[0364] In one aspect, the methods disclosed herein can be used as a primary screening.

[0365] In one aspect, the methods disclosed herein can be used to monitor cancer development.

[0366] In one aspect, the methods disclosed herein can be used to monitor and / or predict cancer risk.

[0367] VII. Identifying or monitoring colorectal cancer

[0368] After processing the data set using the trained algorithm, colorectal cancer can be identified or monitored in the subject. The identification can be based at least in part on the quantitative measure of the sequence reads of the data set of the colorectal cancer-associated genomic locus panel (e.g., the quantitative measure of the RNA transcript or DNA of the colorectal cancer-associated genomic locus).

[0369] Colorectal cancer can be identified in a subject with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The accuracy of identifying colorectal cancer by the trained algorithm can be calculated as the percentage of independently detected samples (e.g., subjects known to have colorectal cancer or subjects with negative clinical test results for colorectal cancer) that are correctly identified or classified as having or not having colorectal cancer.

[0370] Colorectal cancer can be identified in a subject with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The PPV of identifying colorectal cancer using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as having colorectal cancer that correspond to subjects that truly have colorectal cancer.

[0371] Colorectal cancer can be identified in a subject with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The NPV of identifying colorectal cancer using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as not having colorectal cancer that correspond to subjects that do not truly have colorectal cancer.

[0372] Colorectal cancer can be identified in a subject with a clinical sensitivity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater. The clinical sensitivity of identifying colorectal cancer using a trained algorithm can be calculated as the percentage of independent test samples (e.g., subjects known to have colorectal cancer) associated with the presence of colorectal cancer that are correctly identified or classified as having colorectal cancer.

[0373] Colorectal cancer can be identified in a subject with a clinical specificity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater. The clinical specificity of identifying colorectal cancer using a trained algorithm can be calculated as the percentage of independent test samples (e.g., subjects with a negative clinical test result for colorectal cancer) associated with the absence of colorectal cancer that are correctly identified or classified as not having colorectal cancer.

[0374] In some embodiments, the trained algorithm can determine that the subject has a risk of developing colorectal cancer of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater.

[0375] The trained algorithm can determine that the subject has a risk of developing colorectal cancer with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater.

[0376] Upon identifying that the subject has colorectal cancer, the subject can be provided with a therapeutic intervention (e.g., the subject is prescribed or administered an appropriate therapeutic procedure to treat the colorectal cancer). The therapeutic intervention can include a prescribed effective dose of a drug, further testing or evaluation of the colorectal cancer, further monitoring of the colorectal cancer, or a combination thereof. If the subject is currently being treated for the colorectal cancer with one therapeutic procedure, the therapeutic intervention can include a subsequent different therapeutic procedure (e.g., to increase efficacy of treatment due to ineffectiveness of the current therapeutic procedure). The therapeutic intervention can be described by, e.g., “WHO list of priority medical devices for cancer management, WHO Medical device technical series”, World Health Organization, ISBN: 978-92-4-156546-2, Geneva, 2017, the contents of which are incorporated herein by reference. The therapeutic intervention can be described by, e.g., Wolpin et al., “Systemic Treatment of Colorectal Cancer,” Gastroenterology, vol. 134, no. 5, 2008, pages 1296-1310.e1, the contents of which are incorporated herein by reference.

[0377] The therapeutic intervention can include a recommendation that the subject undergo a secondary clinical examination to confirm the diagnosis of the colorectal cancer. This secondary clinical examination can include an imaging examination, a blood examination, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a fecal immunochemical test (FIT), a fecal occult blood test (FOBT), or any combination thereof.

[0378] The quantitative measure of the data set sequence reads on the colorectal cancer-associated genomic locus panel (e.g., the quantitative measure of the RNA transcripts or DNA on the colorectal cancer-associated genomic locus) can be assessed over a period of time to monitor a patient (e.g., a subject having or being treated for colorectal cancer). In this case, the quantitative measure of the patient data set can change over the course of the treatment. For example, the quantitative measure of the patient data set that is at a lower risk of colorectal cancer due to an effective treatment can shift toward the profile or distribution of a healthy subject (e.g., a subject not having colorectal cancer). Conversely, the quantitative measure of the patient data set that is at a higher risk of colorectal cancer due to an ineffective treatment can shift toward the profile or distribution of a subject that is at a higher risk of colorectal cancer or a higher stage or grade of colorectal cancer.

[0379] By monitoring the course of treatment of the subject for colorectal cancer, the subject's colorectal cancer can be monitored. The monitoring can comprise assessing the subject's colorectal cancer at two or more time points. The assessment can be based at least on the quantitative measure of sequence reads of the data set at the colorectal cancer-associated genomic locus panel (e.g., the quantitative measure of RNA transcripts or DNA at the colorectal cancer-associated genomic loci), including the quantitative measure of the colorectal cancer-associated genomic locus panel determined at each of the two or more time points.

[0380] In some embodiments, a difference in the quantitative measure of sequence reads of the data set at the colorectal cancer-associated genomic locus panel (e.g., the quantitative measure of RNA transcripts or DNA at the colorectal cancer-associated genomic loci), including the difference in the quantitative measure of the colorectal cancer-associated genomic locus panel determined between the two or more time points, can be indicative of one or more clinical indications, such as: (i) a diagnosis of colorectal cancer in the subject; (ii) a prognosis of colorectal cancer in the subject; (iii) an increased risk of colorectal cancer in the subject; (iv) a decreased risk of colorectal cancer in the subject; (v) the efficacy of the course of treatment for colorectal cancer in the subject; and (vi) the ineffectiveness of the course of treatment for colorectal cancer in the subject.

[0381] In some embodiments, a difference in the quantitative measure of sequence reads of the data set at the colorectal cancer-associated genomic locus panel (e.g., the quantitative measure of RNA transcripts or DNA at the colorectal cancer-associated genomic loci), including the difference in the quantitative measure of the colorectal cancer-associated genomic locus panel determined between the two or more time points, can be indicative of a diagnosis of colorectal cancer in the subject. For example, if the subject is not detected to have colorectal cancer at an earlier time point, but is detected to have at a later time point, the difference is indicative of a diagnosis of colorectal cancer in the subject. A clinical action or decision can be made based on this indication of a diagnosis of colorectal cancer in the subject, e.g., a new therapeutic intervention is prescribed or administered to the subject. The clinical action or decision can include recommending that the subject undergo a secondary clinical examination to confirm the diagnosis of colorectal cancer. This secondary clinical examination can include an imaging examination, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a fecal cell biology test, a fecal immunochemical test (FIT), a fecal occult blood test (FOBT), or any combination thereof.

[0382] In some embodiments, a difference in quantitative measures of sequence reads of data sets at the colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at the colorectal cancer-associated genomic locus) determined between two or more time points can indicate a prognosis of colorectal cancer in the subject.

[0383] In some embodiments, a difference in quantitative measures of sequence reads of data sets at the colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at the colorectal cancer-associated genomic locus) determined between two or more time points can indicate an increased risk of colorectal cancer in the subject. For example, if the subject is detected to have colorectal cancer at an earlier time point and at a later time point, and if the difference is a positive difference (e.g., the quantitative measures of sequence reads of data sets at the colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at the colorectal cancer-associated genomic locus) are increased from the earlier time point to the later time point), the difference can indicate an increased risk of colorectal cancer in the subject. A clinical action or decision can be made based on this indication of an increased risk of colorectal cancer, for example, prescribing or administering a new therapeutic intervention or switching therapeutic interventions (e.g., ending a current treatment and prescribing or administering a new treatment) for the subject. The clinical action or decision can include recommending a secondary clinical examination for the subject to confirm the increased risk of colorectal cancer. The secondary clinical examination can include an imaging examination, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a

[0384] In some embodiments, a difference in quantitative measures of sequence reads of data sets at a colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at a colorectal cancer-associated genomic locus), including a difference in quantitative measures of the colorectal cancer-associated genomic locus panel determined between two or more time points, can indicate a decrease in the risk of the subject developing colorectal cancer. For example, if the subject is detected to have colorectal cancer at an earlier time point and at a later time point, and if the difference is a negative difference (e.g., quantitative measures of sequence reads of data sets at a colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at a colorectal cancer-associated genomic locus), including quantitative measures of the colorectal cancer-associated genomic locus panel, are decreased from the earlier time point to the later time point), the difference can indicate a decrease in the risk of the subject developing colorectal cancer. A clinical action or decision can be made based on this indication of a decrease in the risk of colorectal cancer, for example, to continue or end a current therapeutic intervention for the subject. The clinical action or decision can include recommending that the subject undergo a secondary clinical examination to confirm the decrease in the risk of developing colorectal cancer. The secondary clinical examination can include an imaging examination, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a non-cellular biological cytology test, a fecal immunochemical test (FIT), a fecal occult blood test (FOBT), or any combination thereof.

[0385] In some embodiments, a difference in quantitative measures of sequence reads of data sets at a colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at a colorectal cancer-associated genomic locus), including a difference in quantitative measures of the colorectal cancer-associated genomic locus panel determined between two or more time points, can indicate the efficacy of a therapeutic process to treat colorectal cancer in the subject. For example, if the subject is detected to have colorectal cancer at an earlier time point, but not at a later time point, the difference can indicate the efficacy of the therapeutic process to treat colorectal cancer in the subject. A clinical action or decision can be made based on this indication of the efficacy of the therapeutic process to treat colorectal cancer in the subject, for example, to continue or end a current therapeutic intervention for the subject. The clinical action or decision can include recommending that the subject undergo a secondary clinical examination to confirm the efficacy of the therapeutic process to treat colorectal cancer in the subject. The secondary clinical examination can include an imaging examination, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a non-cellular biological cytology test, a fecal immunochemical test (FIT), a fecal occult blood test (FOBT), or any combination thereof.

[0386] In some embodiments, a difference in quantitative measures of sequence reads of a dataset at a panel of colorectal cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including a difference in quantitative measures of the panel of colorectal cancer-associated genomic loci determined between two or more time points, can indicate that a treatment process to treat colorectal cancer in the subject is ineffective. For example, if the subject is detected to have colorectal cancer at an earlier time point and at a later time point, and if the difference is a positive or zero difference (e.g., quantitative measures of sequence reads of a dataset at a panel of colorectal cancer-associated genomic loci, including quantitative measures of the panel of colorectal cancer-associated genomic loci, are increasing or remaining at a constant level from the earlier time point to the later time point), the difference can indicate that the treatment process to treat colorectal cancer in the subject is ineffective. A clinical action or decision can be made based on this indication that the treatment process to treat colorectal cancer in the subject is ineffective, e.g., to end the current therapeutic intervention and / or to switch (e.g., prescribe or administer) a new different therapeutic intervention for the subject. The clinical action or decision can include a recommendation for the subject to undergo a secondary clinical examination to confirm the ineffectiveness of the treatment process to treat colorectal cancer in the subject. The secondary clinical examination can include an imaging examination, a blood examination, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell-free biological cytology examination, a fecal immunochemical test (FIT), a fecal occult blood test (FOBT), or any combination thereof.

[0387] VIII. Kits

[0388] The present disclosure provides kits for identifying or monitoring cancer in a subject. The kits can include probes for identifying a quantitative measure of sequence (e.g., indicative of presence, absence, or relative quantity) at each of a plurality of cancer-associated genomic loci in a cell-free biological sample of the subject. The quantitative measure of sequence (e.g., indicative of presence, absence, or relative quantity) at each of the plurality of cancer-associated genomic loci in the cell-free biological sample can be indicative of one or more cancers. The probes can be selective for sequences at the plurality of cancer-associated genomic loci in the cell-free biological sample. The kits can include instructions for processing the cell-free biological sample using the probes to generate a dataset indicative of a quantitative measure of sequence (e.g., indicative of presence, absence, or relative quantity) at each of the plurality of cancer-associated genomic loci in the cell-free biological sample of the subject.

[0389] The probes in the kit can be selective for sequences on a plurality of cancer-associated genomic loci in a cell-free biological sample. The probes in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to the plurality of cancer-associated genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can have sequence complementarity to nucleic acid sequences from one or more of the plurality of cancer-associated genomic loci or genomic regions. The plurality of cancer-associated genomic loci or genomic regions can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more different cancer-associated genomic loci or genomic regions. The plurality of cancer-associated genomic loci or genomic regions can comprise one or more members selected from the regions listed in Tables 1-11.

[0390] The instructions in the kit can include instructions for using probes selective for sequences on a plurality of cancer-associated genomic loci in a cell-free biological sample to assay the cell-free biological sample. The probes can be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity to nucleic acid sequences (e.g., RNA or DNA) from one or more of the plurality of cancer-associated genomic loci. The nucleic acid molecules can be primers or enrichment sequences. The instructions for assaying the cell-free biological sample can include instructions for performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample thereby generating a data set indicative of a quantitative measure of sequences (e.g., indicative of presence, absence, or relative quantity) on each of the plurality of cancer-associated genomic loci in the cell-free biological sample. The quantitative measure of sequences (e.g., indicative of presence, absence, or relative quantity) on each of the plurality of cancer-associated genomic loci in the cell-free biological sample can be indicative of one or more cancers.

[0391] The kit's instructions may include instructions for measuring and interpreting assay readouts that can be quantified at one or more of multiple cancer-associated genomic loci to generate a dataset indicating quantitative measures of sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample (e.g., indicating presence, absence, or relative quantity). For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple cancer-associated genomic loci can generate a dataset indicating quantitative measures of sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample (e.g., indicating presence, absence, or relative quantity). Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values ​​thereof.

[0392] Example

[0393] Example 1: Selection of methylated regions for colorectal cancer detection

[0394] For colorectal cancer, using the systems and methods of this disclosure, 20 highly methylated genomic regions were identified in the tumor, but many of these regions were not methylated in normal tissues. These regions were used as highly specific markers of the presence of tumors with little or no background signal.

[0395] In Table 12, 'Start-End Position' specifies the coordinates of the target region in the construction of the human genome reference sequence hg18. Gene IDs and chromosome fields refer to the gene and chromosome numbers associated with the numbered regions. Examination of these sequences relative to neighboring genes indicates they are found upstream, in 5' promoters, 5' enhancers, introns, exons, distal promoters, coding regions, or intergenic regions, respectively.

[0396] use Cell-free DNA isolation kit (Applied) Cell-free DNA (doped with a unique synthetic double-stranded DNA (dsDNA) fragment for sample tracking) was extracted from 250 microliters (μL) of plasma according to the manufacturer's instructions. Ultra II DNA Library Preparation Kit (New England) Prepare paired-end sequencing libraries, including polymerase chain reaction (PCR) amplification and unique molecular identifiers (UMIs), and use... The NovaSeq 6000 sequencing system sequences at 2x5 l base pairs on multiple S2 or S4 stream cells up to a minimum of 400 million reads (median = 636 million reads).

[0397] probes targeting colorectal cancer

[0398] PCR primer pairs were developed to different regions of the genome that showed extensive methylation in multiple colorectal cancer samples from the TOGA database, but no or little methylation in multiple normal tissues and blood cells (peripheral blood mononuclear cells and others).

[0399] These primers were then used to amplify converted DNA from plasma samples of individuals at risk for colorectal cancer. Sequencing adaptors were ligated to the DNA and next generation sequencing was performed. Sequencing reads were then segregated by region and analyzed using tools such as the BiQ Analyzer HT program.

[0400] The obtained sequencing reads were demultiplexed, adaptor trimmed, and aligned to the human reference genome (GRCh38 with decoys, alt contigs, and HLA contigs) using Burrows Wheeler Aligner (BWA-MEM 0.7.15). PCR duplicate fragments, if present, were removed using fragment endpoints and / or UMIs.

[0401] A cfDNA "profile" was created for each sample by counting the number of fragments that aligned to each putative protein-coding region in the genome. This type of data reveals epigenetic changes in cfDNA that are protected by variable nucleosomes, resulting in observed changes in coverage and methylation of fragments compared to controls.

[0402] A set of functional regions of the human genome, including putative protein-coding gene regions (genomic coordinates range to include introns and exons), were annotated in the sequencing data. Annotations of protein-coding gene regions ("gene" regions) were obtained from the Comprehensive Human Expression Sequence (CHESS) project (vl.O).

[0403] The results obtained are as follows.

[0404] Table 12 provides a set of hypermethylated genomic regions in cell-free nucleic acid samples identified from samples of individuals with colorectal cancer. For each region, an exemplary number of methylated CpG sites in the region is provided as a threshold for distinguishing between healthy individuals and CRC individuals.

[0405] Table 12

[0406]

[0407]

[0408] In the discussion herein, references to genes such as ITGA4, TMEM163, and SFMBT2, for example, can not indicate the gene itself as the gene of interest, but rather the relevant methylation region described in the signature panel.

[0409] A total of 50 regions were found to be hypermethylated in relation to CRC. Not all regions need to be included in the classification model in order to distinguish between healthy individuals and CRC individuals. Therefore, some regions appear to be indicative of various types of cancer in general. Other regions are methylated in subgroups of these, while the rest are specific to cancer. In the context of this assay and the cancer types examined, certain regions can be described as "particularly methylated in colorectal cancer" when training sample sequences in a prediction model, and have a higher weight in the signature. These CRC-related, higher-weighted methylation regions were used in a specific model trained to distinguish between a population of healthy individuals and a population of CRC individuals.

[0410] Example 2: Building and training a classification model for distinguishing a population of colorectal cancer individuals

[0411] Using the system and method of the present disclosure, a machine learning classification model is built and trained using artificial intelligence-based methods to analyze cfDNA data acquired from subjects (to generate a diagnostic output of subjects having colorectal cancer).

[0412] The human plasma samples expected were acquired from 49 patients diagnosed with CRC. In addition, a collection of 92 control samples were acquired from patients who currently have no cancer diagnosis (but can have other co-morbidities or undiagnosed cancers). All samples were de-identified.

[0413] The age, sex, and stage of cancer (when available) of each patient were acquired for each sample. The plasma samples collected from each patient were stored at -80°C and thawed prior to use. Table 13 provides a description of the study cohort, showing the number of healthy and cancer samples (by stage, sex, and age) used for the CRC experiment.

[0414] Table 13

[0415]

[0416]

[0417] Samples were processed and sequenced according to the methods described herein, particularly the methods described in Example 1. The methylation regions in Table 12 were specifically used to determine the methylation CpG status between healthy individuals and individuals with colorectal cancer. For each region listed in column 1 of Table 12, a threshold number of CpG sites shown in column 2 was used to define the methylation fragments for analysis. The remaining fragments were classified as methylated if they had more than the threshold number of CpG sites; otherwise, the fragments were classified as unmethylated. To calculate the raw score for each sample, the counts for each sample were aggregated across regions, which was given by the number of methylated fragments in each sample that overlapped with the regions listed in Table 12. The raw scores for each sample were normalized to account for the coverage differences of each sample. The raw score for each sample was multiplied by a sample-specific scaling factor, which was given by the total number of samples divided by a pre-specified target coverage level. These normalized and scaled methylation ratios were output as the score for each sample. A threshold score was selected based on the desired specificity target from the training set. Based on whether the score of a sample exceeded this threshold, the sample was classified as positive or negative. An ROC curve was generated by considering the ranks of samples with that score or by considering the threshold.

[0418] A machine learning classification model was trained as described above and parameters were selected on a separately proposed set of samples. The machine learning classification model was applied to the samples described in Table 13. The healthy sample with the highest proportion of high methylation fragment counts was selected as the cutoff for classifying new samples as positive or negative. The area under the ROC curve (AUC) was calculated based on the training set described above using the ranks from the normalized high methylation fragment counts. The sensitivity and specificity were calculated with the selected cutoff. The confidence intervals for the sensitivity and specificity were calculated using the Clopper-Pearson confidence interval and the confidence interval for the AUC was calculated using the method described by Fay, M. and Malinovsky, Y., Statistics in Medicine 37(27):3991-4006 (2018), the contents of which are incorporated herein by reference.

[0419] The average area under the curve (AUC) for this method was 0.9488 (0.87-0.98) with an average sensitivity of 70% (0.49-0.87) for the IU samples at 92% specificity (0.86-0.96) Figure 2 ).

[0420] Example 3: Detection of cell-free samples and classification of individuals

[0421] Using the system and methods of the present disclosure, a predictive analysis using artificial intelligence-based methods is used to analyze cfDNA data acquired from a subject to generate a diagnostic output of the subject having colorectal cancer.

[0422] A method is provided herein for predicting an increased risk of developing or having cancer in an asymptomatic patient, wherein a model trained by the signature panel of the process provided in Example 1 is applied to the measured biomarker panel and using clinical factors of age and gender to identify those patients at increased risk of developing or having colorectal cancer. In embodiments, this method and the present classifier model uses input variables of biomarkers measured within a normal clinical range, wherein when the output of the first classifier model is above a calculated threshold based on the number of regional methylation CpG sites, the colorectal cancer classifier model uses an age input variable and measured values from the biomarker panel of the patient to classify the patient into an increased risk category.

[0423] Genes were selected according to Example 1 with the goal of selecting marker genes and CpG sites with strong differential methylation (beta difference, i.e. difference between methylation specific probes and methylation non-specific probes and p-value), predictive power (AUC), and effect on gene expression (p-value from gene expression).

[0424] This selection resulted in the signature panel provided herein, which contains methylation regions that can distinguish between healthy and CRC samples. A first subset of regions contains 20 regions with at least 4 to 18 CpG sites with increased methylation that map to 18 genes (many genes represented by many CpG sites).

[0425] The cfDNA CpG count profile of input cfDNA exhibits an unbiased representation of the methylation signal available in blood, allowing capture of signals directly from the tumor as well as those from non-tumor sources such as the circulating immune system or tumor microenvironment.

[0426] Unsupervised clustering based on these genes shows clear methylation patterns associated with the healthy or CRC phenotype.

[0427] To assess the accuracy of the methylation regions for early detection of CRC, the receiver operating characteristic (ROC) curves and area under the ROC curve (AUC) of the regions in the signature panel were calculated. Figures 3A-3F The ROC results are shown, showing the ability of these differentially methylated regions (DMRs) to detect CRC and distinguish early stage cancer, including stage 1 ( Figure 3A ), stage 2 ( Figure 3B ), stage 3 ( Figure 3C ), stage 4 ( Figure 3D ), missing stage ( Figure 3E ), and all samples ( Figure 3F) patients. A total of 80 gene regions associated with increased methylation were identified. Methylation regions with average methylation levels that gradually increased compared to controls, or that could be used to distinguish early versus late stage CRC. For example, the methylation regions associated in Table 12 had higher CRC detection power [AUC = 0.924 (95% CI: 0.752 to 0.954) for CRC vs. controls].

[0428] As summarized in Table 14, the results demonstrate excellent performance for early cancer detection from blood (e.g., in a set of 13 stage I and II samples).

[0429] Table 14

[0430]

[0431] While preferred embodiments of the application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present application is not intended to be limited by the specific examples provided within the specification. While the application has been described with reference to the above specific embodiment, the description and examples are illustrative of the application and are not intended to limit the scope of the application. Various modifications, changes, and alternatives will become apparent to the skilled artisan upon the reading and understanding of the specification and figures. Furthermore, it should be understood that the application is not limited to the particular description provided herein, but rather is encompassing all modifications, equivalents, and alternatives falling within the scope of the appended claims. The specification and drawings are, accordingly, to be regarded simply as illustrative of the application as defined by the following claims, and are construed in the manner consistent with the scope of the claims.

Claims

1. A methylation signature panel specific for colorectal cancer comprising: methylation genomic regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; CCNA1, chr13: 36431498-36432414; BEND4, chr4: 42150707-42153216; KRBA1, chr7: 149714695-149715338; S1PR1, chr1: 101236505-101237190; PPP1R16B, chr20: 38805341-38807221; IKZF1, chr7: 50304053-50304944; LONRF2, chr2: 100322082-100322599; ZFP82, chr19: 36418330-36418931; FLT3, chr13: 28099881-28100943; FBN1, chr15: 48644595-48646444; and FLI1, chr11: 128693042-128694372, wherein a threshold is determined for the methylation genomic regions, and wherein if the number of methylated CpG sites in the methylation genomic regions in a biological sample exceeds the threshold, then colorectal cancer is indicated.

2. The methylation signature panel of claim 1, wherein the biological sample comprises nucleic acid.

3. The methylation signature panel of claim 2, wherein the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

4. The methylation signature panel of claim 2, wherein the nucleic acid comprises cell-free deoxyribonucleic acid (cfDNA) or cell-free ribonucleic acid (cfRNA).

5. The methylation signature panel of claim 1, wherein the methylation signature panel comprises increased methylation in two or more of the methylation genomic regions.

6. The methylation signature panel of claim 1, wherein the colorectal cancer is selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.

7. A machine learning classifier capable of distinguishing between a population of healthy subjects and a population of subjects having colorectal cancer, comprising: a) a set of measurements representative of the methylation genomic regions of claim 1, wherein the measurements are obtained from methylation sequencing data from healthy subjects and subjects having colorectal cancer; b) wherein the measurements are used to generate a set of features corresponding to properties of the methylation genomic regions, wherein at least one of the properties corresponds to a comparison between the number of methylated CpG sites and a threshold value for the group of the methylation genomic regions, and wherein the set of features is input to a machine learning or statistical model; and c) wherein the machine learning or statistical model provides a feature vector that is used as a classifier capable of distinguishing between a population of healthy subjects and a population of subjects having colorectal cancer.

8. The classifier of claim 7, wherein the set of measurements describes a characteristic of a methylation gene region selected from the group consisting of: percent methylation per base for CpG, CHG, CHH, count or ratio of fragments with different counts or ratios of methylated CpG observed in a region, conversion efficiency, hypomethylated segments, methylation level, overall mean methylation for CPG, CHH, CHG, fragment length, fragment midpoint, number of methylated CpG per fragment, fraction of CpG methylation over total CpG per fragment, fraction of CpG methylation over total CpG per region, fraction of CpG methylation over total CpG in a panel, dinucleotide coverage, coverage uniformity, overall mean CpG coverage, and mean coverage at CpG islands, CGI shores, and CGI shores.

9. A system for detecting colorectal cancer comprising a machine learning model classifier, comprising: a) a computer readable medium comprising a classifier operable to classify a subject as having colorectal cancer or not having colorectal cancer based on a methylation signature panel, wherein the classification is based at least on a comparison between the number of methylated CpG sites and a threshold value for the group of regions of the methylation signature panel, and wherein the methylation signature panel comprises methylation genomic regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; ​ ​ ​ ​ ​ ​ ​ ​ CCNA1, chr13:36431498-36432414; BEND4, chr4:42150707-42153216; KRBA1, chr7:149714695-149715338; S1PR1, chr1:101236505-101237190; PPP1R16B, chr20:38805341-38807221; IKZF1, chr7:50304053-50304944; LONRF2, chr2:100322082-100322599; ZFP82, chr19:36418330-36418931; FLT3, chr13:28099881-28100943; FBN1, chr15:48644595-48646444; and FLI1, chr11:128693042-128694372; and b) one or more processors for executing instructions stored on the computer readable medium.

10. The system of claim 9, comprising the machine learning classifier of claim 6 loaded into a memory of a computer system, a machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having colorectal cancer, and a second subset of the training biological samples identified as not having colorectal cancer.

11. A method of generating a classifier, comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules of a cell-free deoxyribonucleic acid (cfDNA) sample obtained or derived from a subject to produce a plurality of converted nucleic acid molecules; b) contacting the plurality of converted nucleic acid molecules with nucleic acid probes complementary to a pre-identified methylation signature panel of differentially methylated regions to enrich for sequences corresponding to the methylation signature panel, the differentially methylated regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; CCNA1, chr13:36431498-36432414; BEND4, chr4:42150707-42153216; KRBA1, chr7: 149714695-149715338; S1PR1, chr1: 101236505-101237190; PPP1R16B, chr20: 38805341-38807221; IKZF1, chr7: 50304053-50304944; LONRF2, chr2: 100322082-100322599; ZFP82, chr19: 36418330-36418931; FLT3, chr13: 28099881-28100943; FBN1, chr15: 48644595-48646444; and FLI1, chr11: 128693042-128694372; c) comparing the number of methylated CpG sites in the set of differential methylation regions to a threshold, thereby determining a methylation profile of the subject, wherein the threshold is determined for the set of differential methylation regions, and wherein a number of methylated CpG sites in the set of differential methylation regions above the threshold is indicative of colorectal cancer; and d) training a machine learning model to produce the classifier that distinguishes between healthy subjects and subjects having colorectal cancer based at least on the number of methylated CpG sites in the set of differential methylation regions being above the threshold.

12. The method of claim 11, further comprising amplifying the plurality of converted nucleic acids.

13. The method of claim 12, wherein the amplifying comprises polymerase chain reaction (PCR).

14. The method of claim 11, further comprising sequencing the converted nucleic acid molecules at a depth greater than 1000x, thereby generating nucleic acid sequences.

15. The method of claim 14, further comprising sequencing the converted nucleic acid molecules at a depth greater than 2000x, thereby generating nucleic acid sequences.

16. The method of claim 15, further comprising sequencing the converted nucleic acid molecules at a depth greater than 3000x, thereby generating nucleic acid sequences.

17. The method of claim 16, further comprising sequencing the converted nucleic acid molecules at a depth greater than 4000x, thereby generating nucleic acid sequences.

18. The method of claim 17, further comprising sequencing the converted nucleic acid molecules at a depth greater than 5000x, thereby generating nucleic acid sequences.

19. The method of any one of claims 14-18, further comprising aligning the nucleic acid sequences to reference nucleic acid sequences, wherein the reference nucleic acid sequences are at least a portion of a human reference genome.

20. The method of claim 19, wherein the human reference genome is hg18.

21. The method of claim 11, wherein the methylation profile of the subject is indicative of the presence or absence of colorectal cancer in the subject.

22. The method of claim 21, wherein the colorectal cancer is selected from a Stage 1 colorectal cancer, a Stage 2 colorectal cancer, a Stage 3 colorectal cancer, or a Stage 4 colorectal cancer.

23. A system for determining a methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample obtained or derived from a subject, the system comprising: a processor; and a computer readable medium comprising machine executable code that, when executed by the processor, implements a method comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules of the cfDNA sample to produce a plurality of converted nucleic acid molecules; b) contacting the plurality of converted nucleic acid molecules with nucleic acid probes complementary to a pre-identified methylation signature panel of differentially methylated regions to enrich for sequences corresponding to the pre-identified methylation signature panel, the differentially methylated regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; CCNA1, chr13: 36431498-36432414; BEND4, chr4: 42150707-42153216; KRBA1, chr7: 149714695-149715338; S1PR1, chr1: 101236505-101237190; PPP1R16B, chr20: 38805341-38807221; IKZF1, chr7: 50304053-50304944; LONRF2, chr2: 100322082-100322599; ZFP82, chr19: 36418330-36418931; FLT3, chr13: 28099881-28100943; FBN1, chr15: 48644595-48646444; and FLI1, chr11: 128693042-128694372; c) comparing the number of methylated CpG sites in the set of differential methylation regions to a threshold, thereby determining the methylation profile of the subject, wherein the threshold is determined for the set of differential methylation regions, and wherein a number of methylated CpG sites in the set of differential methylation regions exceeding the threshold is indicative of colorectal cancer.

24. The system of claim 23, wherein the method further comprises amplifying the plurality of converted nucleic acid molecules.

25. The system of claim 24, wherein the amplifying comprises polymerase chain reaction (PCR).

26. The system of claim 23, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 1000x, thereby generating nucleic acid sequences.

27. The system of claim 26, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 2000x, thereby generating nucleic acid sequences.

28. The system of claim 27, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 3000x, thereby generating nucleic acid sequences.

29. The system of claim 28, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 4000x, thereby generating nucleic acid sequences.

30. The system of claim 29, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 5000x, thereby generating nucleic acid sequences.

31. The system of claim 26, wherein the method further comprises aligning the nucleic acid sequences to reference nucleic acid sequences, wherein the reference nucleic acid sequences are at least a portion of a human reference genome.

32. The system of claim 31, wherein the human reference genome is hgl8.

33. The system of claim 23, wherein the methylation profile of the subject is indicative of the presence or absence of colorectal cancer in the subject.

34. The system of claim 33, wherein the colorectal cancer is selected from stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, or stage 4 colorectal cancer.

35. A method of generating a classifier, comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules obtained or derived from a biological sample obtained from a subject to produce a plurality of converted nucleic acid molecules; b) contacting the plurality of converted nucleic acid molecules with nucleic acid probes complementary to a pre-identified methylation signature panel of differential methylation regions, to enrich for sequences corresponding to the pre-identified methylation signature panel, the differential methylation regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; CCNA1, chr13: 36431498-36432414; BEND4, chr4: 42150707-42153216; KRBA1, chr7: 149714695-149715338; S1PR1, chr1: 101236505-101237190; PPP1R16B, chr20: 38805341-38807221; IKZF1, chr7: 50304053-50304944; LONRF2, chr2: 100322082-100322599; ZFP82, chr19: 36418330-36418931; FLT3, chr13: 28099881-28100943; FBN1, chr15: 48644595-48646444; and FLI1, chr11: 128693042-128694372; c) determining a threshold value for the set of differentially methylated regions, wherein a number of methylated CpG sites in the set of differentially methylated regions exceeding the threshold value is indicative of colorectal cancer; d) training a machine learning model to produce the classifier, wherein the classifier is trained to be able to distinguish between a healthy subject and a subject having colorectal cancer based on a methylation profile of the subject to provide an output value correlating with the presence of colorectal cancer, thereby detecting the presence or absence of the colorectal cancer in the subject, wherein the output value correlating with the presence of colorectal cancer is determined based at least on a number of methylated CpG sites in the differentially methylated regions exceeding the threshold value.

36. The method of claim 35, wherein the biological sample obtained or derived from the subject is selected from the group consisting of: cell-free deoxyribonucleic acid (cDNA), cell-free ribonucleic acid (cfRNA), a bodily fluid, a stool, a colonic effluent, an isolated blood cell, and combinations thereof.

37. The method of claim 36, wherein the bodily fluid comprises urine, plasma, serum, or whole blood.

38. The method of claim 36, wherein the isolated blood cell comprises a cell isolated from blood.

39. The method of claim 35, further comprising amplifying the plurality of converted nucleic acids.

40. The method of claim 39, wherein the amplifying comprises polymerase chain reaction (PCR).

41. The method of claim 35, further comprising sequencing the converted nucleic acid molecules at a depth greater than 1000x, thereby generating nucleic acid sequences.

42. The method of claim 41, further comprising sequencing the converted nucleic acid molecules at a depth greater than 2000x, thereby generating a nucleic acid sequence.

43. The method of claim 42, further comprising sequencing the converted nucleic acid molecules at a depth greater than 3000x, thereby generating a nucleic acid sequence.

44. The method of claim 43, further comprising sequencing the converted nucleic acid molecules at a depth greater than 4000x, thereby generating a nucleic acid sequence.

45. The method of claim 44, further comprising sequencing the converted nucleic acid molecules at a depth greater than 5000x, thereby generating a nucleic acid sequence.

46. The method of claim 41, further comprising aligning the nucleic acid sequence to a reference nucleic acid sequence, wherein the reference nucleic acid sequence is at least a portion of a human reference genome.

47. The method of claim 46, wherein the human reference genome is hgl8.

48. The method of claim 35, wherein the colorectal cancer is selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.

49. The method of claim 35, wherein the trained machine learning classifier is selected from the group consisting of: a deep learning classifier, a neural network classifier, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and a principal component analysis classifier.

50. A system for detecting the presence or absence of colorectal cancer in a subject, the system comprising: a processor; and a computer readable medium comprising machine executable code that, when executed by the processor, implements a method comprising: a) providing conditions capable of converting unmethylated cytosine to uracil in nucleic acid molecules of a biological sample obtained or derived from the subject to produce a plurality of converted nucleic acid molecules; b) contacting the plurality of converted nucleic acid molecules with nucleic acid probes complementary to a pre-identified panel of methylation signatures of differentially methylated regions comprising: ITGA4, chr2: 181457004-181457950; EMBP1, chr1: 121519076-121519744; TMEM163, chr2: 134718243-134719428; SFMBT2, chr10: 7408046-7408953; ELMO1, chr7: 37448612-37449471; ZNF543, chr19: 57320164-57320845; SFMBT2, chr10: 7410025-7411008; ​ CHST10, chr2: 100417269-100417795; ELMO1, chr7: 37447852-37448217; CCNA1, chr13: 36431498-36432414; BEND4, chr4: 42150707-42153216; KRBA1, chr7: 149714695-149715338; S1PR1, chr1: 101236505-101237190; PPP1R16B, chr20: 38805341-38807221; IKZF1, chr7: 50304053-50304944; LONRF2, chr2: 100322082-100322599; ZFP82, chr19: 36418330-36418931; FLT3, chr13: 28099881-28100943; FBN1, chr15: 48644595-48646444; and FLI1, chr11: 128693042-128694372; c) comparing the number of methylated CpG sites in the set of differential methylation regions to a threshold, thereby determining a methylation profile of the subject, wherein the threshold is determined for the set of differential methylation regions, and wherein an excess of methylated CpG sites in the set of differential methylation regions over the threshold is indicative of colorectal cancer; and d) applying a trained machine learning classifier to the methylation profile of the subject, wherein the trained machine learning classifier is trained to be able to distinguish between a healthy subject and a subject having colorectal cancer, to provide an output value associated with the presence of colorectal cancer, thereby detecting the presence or absence of the colorectal cancer in the subject, wherein the output value associated with the presence of colorectal cancer is at least partially determined based on an excess of methylated CpG sites in the set of differential methylation regions over the threshold.

51. The system of claim 50, wherein the biological sample obtained or derived from the subject is selected from the group consisting of: cell-free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), a bodily fluid, a stool, a colonic effluent, an isolated blood cell, and combinations thereof.

52. The system of claim 51, wherein the bodily fluid comprises urine, plasma, serum, or whole blood.

53. The system of claim 51, wherein the isolated blood cell comprises a cell isolated from blood.

54. The system of claim 50, wherein the method further comprises amplifying the plurality of transformed nucleic acid molecules.

55. The system of claim 54, wherein the amplifying comprises polymerase chain reaction (PCR).

56. The system of claim 50, wherein the method further comprises sequencing the transformed nucleic acid molecules at a depth greater than 1000x, thereby generating nucleic acid sequences.

57. The system of claim 56, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 2000x, thereby generating nucleic acid sequences.

58. The system of claim 57, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 3000x, thereby generating nucleic acid sequences.

59. The system of claim 58, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 4000x, thereby generating nucleic acid sequences.

60. The system of claim 59, wherein the method further comprises sequencing the converted nucleic acid molecules at a depth greater than 5000x, thereby generating nucleic acid sequences.

61. The system of claim 56, further comprising aligning the nucleic acid sequences to reference nucleic acid sequences, wherein the reference nucleic acid sequences are at least a portion of a human reference genome.

62. The system of claim 61, wherein the human reference genome is hg18.

63. The system of claim 50, wherein the method further comprises administering a treatment for the colorectal cancer to the subject based on detecting the presence of the colorectal cancer in the subject.

64. The system of claim 50, wherein the colorectal cancer is selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.

65. The system of claim 50, wherein the trained machine learning classifier is selected from the group consisting of: a deep learning classifier, a neural network classifier, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and a principal component analysis classifier.

Citation Information

Patent Citations

  • Method of detection of methylated nucleic acid using agents which modify unmethylated cytosine and distinguishing modified methylated and non-methylated nucleic acids

    US5786146A

  • Use of microarrays for genomic representation selection

    US20080194413A1

  • Machine learning implementation for multi-analyte assay of biological samples

    WO2019200410A1