Method and system for detecting colorectal cancer by nucleic acid methylation analysis
By analyzing genomic methylation status using a methylation signature panel and machine learning classifier, this method addresses the trade-off between false positives and false negatives in existing tools, achieving highly sensitive and specific detection of colorectal cancer, and is suitable for early screening and monitoring.
Patent Information
- Application Number
- CN202511382220.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-31
- Filing Date
- 2021-03-29
- Publication Date
- 2026-03-06
AI Technical Summary
Existing cancer screening tools present a trade-off between false positive and false negative results in colorectal cancer detection, and the sensitivity of circulating tumor DNA testing is limited, making it difficult to detect early signs of colorectal cancer with low tumor burden.
Using a methylation signature panel and a machine learning classifier, this study distinguishes healthy individuals from those with colonic cell proliferation disorder by analyzing the genomic methylation status in individual biological samples, and uses methylation regions as biomarkers for detection.
It improves the sensitivity and specificity of colorectal cancer detection, enabling early identification of colorectal cancer with low tumor burden, providing a more accurate screening tool, and is suitable for initial screening and disease progression monitoring in high-risk populations.
Smart Images

Figure CN121610573A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202180039398.8, filed on March 29, 2021, entitled "Method and System for Detecting Colorectal Cancer by Nucleic Acid Methylation Analysis" (the corresponding PCT application was filed on March 29, 2021, with application number PCT / US2021 / 024604).
[0002] Cross-references to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 002,878, filed March 31, 2020, the contents of which are hereby incorporated by reference in their entirety. Background Technology
[0004] This disclosure generally relates to cancer detection and disease surveillance. More specifically, the field involves cancer-related DNA methylation detection and disease surveillance for early colorectal cancer (CRC). Over the past few decades, cancer screening and surveillance have helped improve outcomes because early detection leads to better results, allowing cancer to be eliminated before it spreads. For example, in the case of CRC, colonoscopy can play a role in improving early diagnosis. Unfortunately, some challenges can arise due to patients' lack of adherence to recommended screening patterns.
[0005] A major problem with any screening tool can be the trade-off between false positives and false negatives (or specificity and sensitivity), the former leading to unnecessary investigations and the latter to ineffectiveness. Ideally, a test should have a high positive predictive value (PPV), minimizing unnecessary investigations while still detecting the vast majority of cancers. Another key factor is the so-called "detection sensitivity," which is also the lower limit of detection based on tumor size, distinguishing it from test sensitivity. Unfortunately, waiting for a tumor to grow large enough to release the levels of circulating tumor markers necessary for detection may conflict with the requirement for early detection, which aims to treat tumors at the most effective stage of treatment. Therefore, effective blood-based screening for early CRC based on circulating analytes is needed.
[0006] Circulating tumor DNA (CBTC) detection is increasingly recognized as a viable “liquid biopsy,” allowing for non-invasive tumor detection and information gathering. In some cases, these techniques have been applied to colorectal, breast, and prostate cancers for the identification of tumor-specific mutations. However, the sensitivity of these techniques may be limited by the presence of high background levels of normal (e.g., non-tumor-derived) DNA in circulation.
[0007] Detection of tumor-specific methylation in the blood offers significant advantages over mutation detection. Many single or polymethylation biomarkers can be evaluated in cancers including lung, colon, and breast cancer. These may have low sensitivity because they may not be prevalent enough in tumors.
[0008] More sensitive and specific screening tools are still needed to detect early or low-tumor-burden colorectal cancer tumor signals in recurrence and to conduct initial screening in high-risk groups. Summary of the Invention
[0009] This disclosure provides methods and systems for analyzing gene methylation profiles related to colorectal cancer detection and disease progression.
[0010] On the one hand, this disclosure provides a methylation signature panel specific to colonic proliferative disorders, comprising: one or more methylated genomic regions selected from Table 11, wherein the one or more regions are more methylated in biological samples from individuals with colonic proliferative disorders or subtypes of colonic proliferative disorders, and less methylated in normal tissues and normal blood cells from individuals without colonic proliferative disorders.
[0011] In some implementations, the biological sample is nucleic acid, DNA, ribonucleic acid (RNA), or cell-free nucleic acid (e.g., cfDNA or cfRNA).
[0012] In some implementation schemes, genomic regions are divided into non-coding regions, coding regions, non-transcribed regions, or regulatory regions.
[0013] In some implementations, the signature panel includes increased methylation in two or more genomic regions selected from Table 11.
[0014] In some implementations, the biological sample obtained from the subject is selected from: cell-free DNA, cell-free RNA, body fluids, feces, colon excretions, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0015] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0016] In some implementations, the colonic cell proliferative disorder is selected from stage 1, stage 2, stage 3, or stage 4 colorectal cancer.
[0017] In some implementations, the signature panel includes two or more methylated genomic regions from Tables 1-11, three or more methylated genomic regions from Tables 1-11, four or more methylated genomic regions from Tables 1-11, five or more methylated genomic regions from Tables 1-11, six or more methylated genomic regions from Tables 1-11, seven or more methylated genomic regions from Tables 1-11, eight or more methylated genomic regions from Tables 1-11, nine or more methylated genomic regions from Tables 1-11, ten or more methylated genomic regions from Tables 1-11, eleven or more methylated genomic regions from Tables 1-11, twelve or more methylated genomic regions from Tables 1-11, or thirteen or more methylated genomic regions from Tables 1-11.
[0018] In some implementations, the signature panel includes methylated genomic regions in colorectal cancer, including methylated regions selected from one or more of the following genomic regions: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, and ZNF543.
[0019] In some implementations, the methylated regions in colorectal cancer include methylated regions in the ITGA4 and EMBP1 genomic regions.
[0020] In some implementations, the methylated regions in colorectal cancer include methylated regions selected from one or more of the following genomic regions: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.
[0021] In some implementations, the signature panel includes methylated genomic regions selected from Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11.
[0022] On the other hand, this disclosure provides a methylation signature panel specific to colonic proliferative disorders, comprising two or more methylated genomic regions in Tables 1-11, wherein the two or more regions are more methylated in biological samples from individuals with colonic proliferative disorders or subtypes of colonic proliferative disorders, and less methylated in normal tissues and normal blood cells from individuals without colonic proliferative disorders.
[0023] In some implementations, the biological sample is nucleic acid, DNA, ribonucleic acid (RNA), or cell-free nucleic acid (cfDNA or cfRNA).
[0024] In some implementation schemes, genomic regions are divided into non-coding regions, coding regions, non-transcribed regions, or regulatory regions.
[0025] In some implementations, the signature panel includes methylation increases in 6 or more or 12 or more genomic regions listed in Tables 1-11.
[0026] In some implementations, the biological sample obtained from the subject is selected from: cell-free DNA, cell-free RNA, body fluids, feces, colon excretions, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0027] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0028] In some implementations, the colonic cell proliferative disorder is selected from stage 1, stage 2, stage 3, or stage 4 colorectal cancer.
[0029] In some implementations, the signature panel includes three or more methylated genomic regions from Tables 1-11, four or more methylated genomic regions from Tables 1-11, five or more methylated genomic regions from Tables 1-11, six or more methylated genomic regions from Tables 1-11, seven or more methylated genomic regions from Tables 1-11, eight or more methylated genomic regions from Tables 1-11, nine or more methylated genomic regions from Tables 1-11, ten or more methylated genomic regions from Tables 1-11, eleven or more methylated genomic regions from Tables 1-11, twelve or more methylated genomic regions from Tables 1-11, or thirteen or more methylated genomic regions from Tables 1-11.
[0030] In some implementations, the signature panel includes methylated genomic regions in colorectal cancer, including methylated regions selected from one or more of the following genomic regions: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, and ZNF543.
[0031] In some implementations, the methylated regions in colorectal cancer include methylated regions in the ITGA4 and EMBP1 genomic regions.
[0032] In some implementations, the methylated regions in colorectal cancer include methylated regions selected from one or more of the following genomic regions: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.
[0033] In some implementations, the signature panel includes methylated regions selected from Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11.
[0034] On the other hand, this disclosure provides a classifier (e.g., a machine learning classifier) capable of distinguishing between a group of healthy individuals and a group of individuals with colonic proliferative disease, comprising: a) a set of measurements representing differentially methylated genomic regions, wherein the measurements are obtained from methylation sequencing data from healthy subjects and subjects with colonic proliferative disease; b) wherein the measurements are used to generate a feature set corresponding to the characteristics of the differentially methylated genomic regions, and the features are input into a machine learning or statistical model; and c) wherein the model provides a feature vector used as a classifier capable of distinguishing between a group of healthy individuals and individuals with colonic proliferative disease.
[0035] In some embodiments, the set of measurements describes characteristics of methylated regions selected from: the percentage of base-by-base methylation of CpG, CHG, and CHH; the count or ratio of fragments with different counts or ratios of methylated CpG observed in the region; conversion efficiency (100 - average methylation percentage of CHH); low-methylated segments; methylation level (overall average methylation of CpG, CHH, and CHG); fragment length; fragment midpoint; and in one or more genomic regions such as chrM, LINE1, or ALU. The methylation level in the domain), the number of methylated CpGs in each fragment, the percentage of CpG methylation in each fragment relative to the total CpG, the percentage of CpG methylation in each region relative to the total CpG, the percentage of CpG methylation in the panel relative to the total CpG, dinucleotide coverage (normalized dinucleotide coverage), coverage uniformity (unique CpG sites under 1x and 10x average genome coverage (for S4 run)), overall average CpG coverage (depth), and average coverage at CpG islands, CGI racks, and CGI shores.
[0036] In some implementations, the machine learning model includes a classifier loaded into the memory of a computer system, a machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having colonic proliferative disease, and a second subset of the training biological samples identified as not having colonic proliferative disease.
[0037] In some implementations, the classifier is provided in a system for detecting colonic proliferative disorder, the system comprising: a) a computer-readable medium including the classifier, the classifier being operable to classify objects as having or not having colonic proliferative disorder based on a methylation signature panel; and b) one or more processors for executing instructions stored on the computer-readable medium.
[0038] In some implementations, the system includes a classification loop configured to use a machine learning classifier selected from: deep learning classifiers, neural network classifiers, linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, linear kernel support vector machine classifiers, first-order or second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic net algorithm classifiers, sequence minimum optimization algorithm classifiers, Naive Bayes algorithm classifiers, and principal component analysis classifiers.
[0039] In some implementations, a computer-readable medium is a non-transitory computer-readable medium that includes machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0040] In some implementations, the system includes one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described herein.
[0041] In another aspect, this disclosure provides a method for determining the methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample from an individual, comprising: a) providing conditions capable of converting unmethylated cytosine to uracil in nucleic acid molecules of the cfDNA sample to produce a plurality of transformed nucleic acids; b) contacting the plurality of transformed nucleic acids with a nucleic acid probe complementary to a pre-identified methylation signature panel selected from at least two differentially methylated regions in Tables 1-11 to enrich sequences corresponding to the signature panel; c) determining the nucleic acid sequences of the plurality of transformed nucleic acid molecules; and d) aligning the nucleic acid sequences of the plurality of transformed nucleic acid molecules with a reference nucleic acid sequence to determine the methylation profile of the individual.
[0042] In some embodiments, the nucleic acid sequencing library is prepared prior to amplification. In some embodiments, the method further includes amplifying multiple transformed nucleic acids. In some embodiments, the amplification includes polymerase chain reaction (PCR). In some embodiments, the method further includes determining the nucleic acid sequence of the transformed nucleic acid molecules at depths greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x. In some embodiments, the reference nucleic acid sequence is at least a portion of the human reference genome. In some embodiments, the human reference genome is hg18.
[0043] In some implementations, the methylation profile is associated with colonic cell proliferative disorder and provides a classification of subjects with colonic cell proliferative disorder.
[0044] In some implementations, a nucleic acid aptamer containing a unique molecular identifier is ligated to the unconverted nucleic acid in the cfDNA sample prior to a).
[0045] In some implementations, nucleic acid molecules are subjected to conditions of cytosine to uracil conversion using chemical methods, enzymatic methods, or combinations thereof.
[0046] In some implementations, the cfDNA in the biological sample is treated with a reagent selected from the following: bisulfite, hydrogen sulfite, disulfide, and combinations thereof.
[0047] In some implementations, the biological sample obtained from the subject is selected from: cell-free DNA, cell-free RNA, body fluids, feces, colon excretions, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0048] In some embodiments, the method includes comparing a methylation signature panel measured from a subject with a database of methylation signature panels measured from normal subjects, wherein the database is stored in a computer system; determining an increased risk of the subject having colonic proliferative disease by measuring a change in the methylation state of the methylation signature panel of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, or at least 20% compared to the methylation state from normal subjects.
[0049] In some embodiments, the pre-identification methylation signature panel includes three or more methylated genomic regions from Tables 1-11, four or more methylated genomic regions from Tables 1-11, five or more methylated genomic regions from Tables 1-11, six or more methylated genomic regions from Tables 1-11, seven or more methylated genomic regions from Tables 1-11, eight or more methylated genomic regions from Tables 1-11, nine or more methylated genomic regions from Tables 1-11, ten or more methylated genomic regions from Tables 1-11, eleven or more methylated genomic regions from Tables 1-11, twelve or more methylated genomic regions from Tables 1-11, or thirteen or more methylated genomic regions from Tables 1-11. In some embodiments, the pre-identification methylation signature panel includes one or more methylated genomic regions from Table 11, two or more methylated genomic regions from Table 11, or three methylated genomic regions from Table 11. In some embodiments, the methylation profile indicates the presence or absence of colonic proliferative disease in an individual.
[0050] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0051] In some implementations, the colonic cell proliferative disorder is selected from stage 1, stage 2, stage 3, and stage 4 colorectal cancer.
[0052] In another aspect, this disclosure provides a method for detecting the presence or absence of colonic proliferative disorder (CPD) in a subject, comprising: a) providing conditions for converting unmethylated cytosine to uracil in nucleic acid molecules obtained from or derived from the subject to produce a plurality of transformed nucleic acids; b) contacting the plurality of transformed nucleic acids with a nucleic acid probe complementary to a pre-identified methylation signature panel selected from at least two differentially methylated regions in Tables 1-11 to enrich sequences corresponding to the signature panel; c) determining the nucleic acid sequences of the plurality of transformed nucleic acid molecules; d) aligning the nucleic acid sequences of the plurality of transformed nucleic acid molecules with a reference nucleic acid sequence to determine the methylation profile of the individual; and e) applying a trained machine learning model to the methylation profile, wherein the trained machine learning model is trained to distinguish between healthy individuals and individuals with CPD to provide output values associated with the presence of CPD, thereby detecting the presence or absence of CPD in the subject.
[0053] In some embodiments, the nucleic acid sequencing library is prepared prior to amplification. In some embodiments, the method further includes amplifying multiple transformed nucleic acids. In some embodiments, the amplification includes polymerase chain reaction (PCR). In some embodiments, the method further includes determining the nucleic acid sequence of the transformed nucleic acid molecules at depths greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x. In some embodiments, the reference nucleic acid sequence is at least a portion of the human reference genome. In some embodiments, the human reference genome is hg18.
[0054] In some implementations, the biological sample obtained from the subject is selected from: cell-free DNA, cell-free RNA, body fluids, feces, colon excretions, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0055] In some embodiments, the method includes comparing a methylation signature panel measured from a subject with a database of methylation signature panels measured from normal subjects, wherein the database is stored in a computer system; determining an increased risk of the subject having colonic proliferative disease by measuring a change in the methylation state of the methylation signature panel of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, or at least 20% compared to the methylation state from normal subjects.
[0056] In some embodiments, the pre-identification methylation signature panel includes three or more methylated genomic regions from Tables 1-11, four or more methylated genomic regions from Tables 1-11, five or more methylated genomic regions from Tables 1-11, six or more methylated genomic regions from Tables 1-11, seven or more methylated genomic regions from Tables 1-11, eight or more methylated genomic regions from Tables 1-11, nine or more methylated genomic regions from Tables 1-11, ten or more methylated genomic regions from Tables 1-11, eleven or more methylated genomic regions from Tables 1-11, twelve or more methylated genomic regions from Tables 1-11, or thirteen or more methylated genomic regions from Tables 1-11. In some embodiments, the pre-identification methylation signature panel includes one or more methylated genomic regions from Table 11, two or more methylated genomic regions from Table 11, or three methylated genomic regions from Table 11. In some embodiments, the methylation profile indicates the presence or absence of colonic proliferative disorders in an individual. In some implementations, the method further includes administering treatment for colonic cell proliferation disorder to the individual based on the detection of the presence of colonic cell proliferation disorder in the individual.
[0057] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0058] In some implementations, the trained machine learning classifier is selected from: deep learning classifiers, neural network classifiers, linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, linear kernel support vector machine classifiers, first-order or second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic net algorithm classifiers, sequence minimum optimization algorithm classifiers, Naive Bayes algorithm classifiers, and principal component analysis classifiers.
[0059] In some implementations, the colonic cell proliferative disorder is selected from stage 1, stage 2, stage 3, and stage 4 colorectal cancer.
[0060] On the other hand, this disclosure provides a method for monitoring minimal residual disease in subjects who have previously received treatment for a disease, comprising: determining the methylation spectrum described herein as a baseline methylation state, and repeating the analysis to determine the methylation spectrum at one or more predetermined time points, wherein changes compared to the baseline indicate changes in the subject's minimal residual disease status at the baseline.
[0061] In some implementations, minimal residual disease is selected from response to treatment, tumor burden, postoperative residual tumor, recurrence, secondary screening, primary screening, and cancer progression.
[0062] On the other hand, a method for determining the response to treatment is provided.
[0063] On the other hand, a method for monitoring tumor burden is provided.
[0064] On the other hand, a method for detecting residual tumors after surgery is provided.
[0065] On the other hand, a method for detecting recurrence is provided.
[0066] On the other hand, a method for secondary screening is provided.
[0067] On the other hand, a method is provided for use as a screening.
[0068] On the other hand, a method for monitoring cancer progression is provided.
[0069] In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 80% sensitivity. In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 90% sensitivity. In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 95% sensitivity. In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 70% positive predictive value (PPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 80% positive predictive value (PPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 90% positive predictive value (PPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 95% positive predictive value (PPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer with at least about 99% positive predictive value (PPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer at at least about 80% negative predictive value (NPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer at at least about 90% negative predictive value (NPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer at at least about 95% negative predictive value (NPV). In some embodiments, the dataset indicates the presence or susceptibility of colorectal cancer at at least about 99% negative predictive value (NPV). In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in subjects with an area under the curve (AUC) of at least about 0.90. In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in subjects with an area under the curve (AUC) of at least about 0.95. In some embodiments, the trained algorithm determines the presence or susceptibility of colorectal cancer in subjects with an area under the curve (AUC) of at least about 0.99.
[0070] In some embodiments, the method further includes displaying the report on a graphical user interface of the user's electronic device. In some embodiments, the user is an object, an individual, or a patient.
[0071] In some implementations, the method further includes determining the likelihood of the presence or susceptibility to colorectal cancer in a subject, individual, or patient. For example, the likelihood may be a probability value between 0% and 100%.
[0072] In some implementations, the trained algorithm (e.g., a machine learning model or classifier) includes a supervised machine learning algorithm. In some implementations, supervised machine learning algorithms include deep learning algorithms, support vector machines (SVMs), neural networks, or random forests.
[0073] In some implementations, the method further includes providing the subject with a therapeutic intervention at least in part based on methylation profiling or analysis, such as a therapeutic intervention for treating colorectal cancer patients (e.g., chemotherapy, radiotherapy, immunotherapy, or surgery).
[0074] In some embodiments, the method further includes monitoring the presence or susceptibility to colorectal cancer, wherein the monitoring includes assessing the presence or susceptibility to colorectal cancer in the subject at multiple time points, wherein the assessment is based at least on the presence or susceptibility to colorectal cancer determined at each of the multiple time points.
[0075] In some implementations, the differential indications for the assessment of the presence or susceptibility to colorectal cancer in subjects at multiple time points are selected from one or more of the following clinical indications: (i) diagnosis of the presence or susceptibility to colorectal cancer in subjects, (ii) prognosis of the presence or susceptibility to colorectal cancer in subjects, and (iii) efficacy or ineffectiveness of treatment in the presence or susceptibility to colorectal cancer in subjects.
[0076] In some implementations, the method further includes stratifying the colorectal cancer of the subject by using a trained algorithm to determine the colorectal cancer subtype of the subject from multiple different colorectal cancer subtypes or stages.
[0077] Another aspect of this disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0078] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.
[0079] Further aspects and advantages of this disclosure will readily become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of this disclosure are shown and described. As will be understood, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious ways, all without departing from this disclosure. Therefore, the drawings and description are to be regarded in an illustrative rather than restrictive manner.
[0080] Incorporation
[0081] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference as if each individual publication, patent, or patent application were expressly and individually incorporated by reference. Where a publication, patent, or patent application incorporated by reference contradicts the disclosure contained in this specification, this specification is intended to supersede and / or give precedence to any such contradictory material. Attached Figure Description
[0082] Embodiments of this disclosure will now be described by way of example only with reference to the accompanying drawings. The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments utilizing the principles of the invention, along with the accompanying drawings (also referred to herein as “Figures”), in which:
[0083] Figure 1 Schematic diagrams are provided illustrating the programming of computer systems or other configuration of machine learning models and classifiers to implement the methods presented in this paper.
[0084] Figure 2 The area under the curve (AUC) curves for 4x cross-validation of the models trained on the regions in Table 1 are provided.
[0085] Figures 3A-3F A series of area under the curve (AUC) curves are provided for samples at different stages of CRC training on the classification model. Figures 3A-3F The ROC results are shown, demonstrating the ability of these differentially methylated regions (DMRs) to detect CRC and differentiate early-stage cancers, including those with stage 1 ( Figure 3A Phase 2 Figure 3B Phase 3 Figure 3C ), 4th phase ( Figure 3D ), missing phase ( Figure 3E ) and all samples ( Figure 3F (patients) Detailed Implementation
[0086] Although various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives may be employed to the embodiments of the invention described herein.
[0087] This disclosure generally relates to cancer detection and disease surveillance. More specifically, the field involves cancer-related DNA methylation detection and disease surveillance for early colorectal cancer. Over the past few decades, cancer screening and surveillance have helped improve outcomes because early detection leads to better results, allowing cancer to be eliminated before it spreads. In the case of colorectal cancer, for example, colonoscopy can play a role in improving early diagnosis. Unfortunately, some challenges can arise due to patients' lack of adherence to recommended screening schedules.
[0088] A major problem with any screening tool can be the trade-off between false positives and false negatives (or specificity and sensitivity), the former leading to unnecessary investigations and the latter to ineffectiveness. Ideally, a test should have a high positive predictive value (PPV), minimizing unnecessary investigations while detecting the vast majority of cancers. Another key factor is the so-called "detection sensitivity," which is also the lower limit of detection based on tumor size, distinguishing it from test sensitivity. Unfortunately, waiting for a tumor to grow large enough to release the levels of circulating tumor markers necessary for detection may conflict with the requirement for early detection, which aims to treat tumors at the most effective stage of treatment. Therefore, effective blood-based screening for early colorectal cancer based on circulating analytes is needed.
[0089] Circulating tumor DNA (CBTC) detection is increasingly recognized as a viable “liquid biopsy,” allowing for non-invasive tumor detection and information gathering. In some cases, these techniques have been applied to colorectal, breast, and prostate cancers for the identification of tumor-specific mutations. However, the sensitivity of these techniques may be limited by the presence of high background levels of normal (e.g., non-tumor-derived) DNA in circulation.
[0090] Detection of tumor-specific methylation in the blood offers significant advantages over mutation detection. Many single or polymethylation biomarkers can be evaluated in cancers including lung, colon, and breast cancer. These may have low sensitivity because they may not be prevalent enough in tumors.
[0091] More sensitive and specific screening tools are still needed to detect early or low-tumor-burden colorectal cancer tumor signals in recurrence and to conduct initial screening in high-risk groups.
[0092] This disclosure provides methods and systems for analyzing gene methylation profiles related to colorectal cancer detection and disease progression.
[0093] On the one hand, this disclosure provides methods for using methylated region panels suitable for analyzing methylation within regions or genes; on the other hand, it provides new uses for said regions, genes, and gene products, as well as methods, assays, and kits relating to the detection, differentiation, and distinction of colonic proliferative disorders. The methods and nucleic acids provided herein can be used to analyze colonic proliferative disorders selected from adenocarcinoma, adenoma, polyp, squamous cell carcinoma, carcinoid tumor, sarcoma, and lymphoma.
[0094] In some embodiments, the method includes using one or more genes selected from methylated regions as biomarkers for the differentiation, detection, and distinction of colonic proliferative disorders. The use of these genes can be enabled by analyzing the methylation status of one or more genes selected from the methylated regions described herein, as well as their promoters or regulatory elements.
[0095] The methods and systems disclosed herein may include analyzing the methylation status of CpG dinucleotides in one or more genomic sequences based on the methylation regions and complementary sequences described herein.
[0096] I. Definition
[0097] Unless the context clearly indicates otherwise, as used in the specification and claims, the singular forms "a / an" and "the" include a plural of indicators. For example, the term "nucleic acid" includes a plurality of nucleic acids, including mixtures thereof.
[0098] As used herein, the term "object" generally refers to an entity or medium that has testable or detectable genetic information. An object can be a person, an individual, or a patient. An object can be a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, apes, farm animals, treadmills, rodents, and pets. An object can be a person who has cancer or is suspected of having cancer. An object can exhibit symptoms indicative of its health or physiological state or condition, such as cancer or other diseases, ailments, or symptoms. Alternatively, an object may be asymptomatic in relation to such health or physiological state or condition.
[0099] As used herein, the term "sample" generally refers to a biological sample obtained from or derived from one or more objects. A biological sample may be a cell-free biological sample or a substantially cell-free biological sample, or it may be processed or fractionated to produce a cell-free biological sample. For example, cell-free biological samples may include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and their derivatives. EDTA collection tubes and cell-free RNA collection tubes (e.g., [missing information]) may be used. ) or cell-free DNA collection tubes (e.g. Cell-free biological samples are obtained or derived from an object. Cell-free biological samples can be derived from whole blood samples by fractionation (e.g., centrifugation of cellular and cell-free components). Biological samples or their derivatives may contain cells. For example, a biological sample may be a blood sample or its derivative (e.g., blood collected via a collection tube or blood droplet).
[0100] As used herein, the term "nucleic acid" generally refers to a polymer of nucleotides of any length, whether deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nucleic acids can have any three-dimensional structure and can perform any known or unknown function. Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coding or non-coding regions of genes or gene segments, loci as defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched-chain nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be conferred before or after nucleic acid assembly. The nucleotide sequence of a nucleic acid can be interrupted by non-nucleotide components. Nucleic acids can be further modified after polymerization, such as through conjugation or binding to reporter factors.
[0101] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule in a population of starting nucleic acid molecules, the presence, number, and / or variation of the nucleotide sequence of which needs to be determined. A target nucleic acid can be any type of nucleic acid, including DNA, RNA, and their analogues. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.
[0102] As used herein, the terms “amplifying” and “amplification” generally refer to increasing the size or number of nucleic acid molecules. Nucleic acid molecules can be single-stranded or double-stranded. Amplification may include generating one or more copies of the nucleic acid molecule, or “amplification product.” Amplification can be performed, for example, by extension (e.g., primer extension) or ligation. Amplification may include performing a primer extension reaction to generate a strand complementary to the single-stranded nucleic acid molecule, and in some cases, generating one or more copies of that strand and / or the single-stranded nucleic acid molecule. The term “DNA amplification” generally refers to generating one or more copies of a DNA molecule, or “amplified DNA product.” The term “reverse transcription amplification” generally refers to the generation of deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template by the action of reverse transcriptase.
[0103] As used herein, the term "cell-free nucleic acid (cfNA)" generally refers to nucleic acids (such as cell-free RNA ("cfRNA") or cell-free DNA ("cfDNA")) that are not contained in cells in a biological sample. cfDNA can circulate freely in bodily fluids such as in the bloodstream.
[0104] As used herein, the term "cell-free sample" generally refers to a biological sample that is substantially devoid of intact cells. This can be derived from biological samples that are substantially devoid of cells themselves, or from samples in which cells have been removed. Examples of cell-free samples include those derived from blood, such as serum or plasma; urine; or samples derived from other sources, such as semen, sputum, feces, catheter exudate, lymph, or recovered lavage fluid.
[0105] As used in this article, the term "circulating tumor DNA" generally refers to cfDNA derived from tumors.
[0106] As used herein, the term "genomic region" generally refers to the identifiable regions of nucleic acids, identified by their location within the chromosome. In some instances, a genomic region is referred to by a gene name and encompasses both coding and non-coding regions associated with the physical regions of nucleic acids. As used herein, a gene contains coding regions (exons), non-coding regions (introns), transcriptional control regions or other regulatory regions, and promoters. In another instance, a genomic region may incorporate introns or exons, or intron / exon boundaries, within a named gene.
[0107] As used herein, the term “CpG island” generally refers to a contiguous region of genomic DNA that meets the following criteria: (1) a frequency of CpG dinucleotides corresponding to an “observed / expected ratio” greater than about 0.6; and (2) a “GC content” greater than about 0.5. CpG islands are typically (but not always) between 0.2 and 3 kilobases (kb) in length and have a high frequency of CpG sites. CpG islands are found at or near the promoters of about 40% of mammalian genes. CpG islands are also found outside of mammalian genes. In some instances, CpG islands are found in exons, introns, promoters, enhancers, repressors, and transcriptional regulatory elements. CpG islands may tend to appear upstream of so-called “housekeeping genes.” The CpG dinucleotide content of CpG islands is said to be at least about 60% of the statistically expected content. The presence of CpG islands at or upstream of the 5' end of a gene may reflect their role in transcriptional regulation, and methylation of CpG sites within a gene promoter may lead to silencing. Conversely, the silencing of tumor suppressors caused by methylation is a hallmark of many human cancers.
[0108] As used in this article, the term "CpG shore" generally refers to a short-distance region extending outward from a CpG island, where methylation may also occur. CpG shores can be found in regions approximately 0 to 2 kb upstream and downstream of CpG islands.
[0109] As used herein, the term “CpG rack” generally refers to a short-distance region extending from the CpG shore, where methylation may also occur. CpG racks are generally found in regions approximately 2 kb to 4 kb upstream and downstream of CpG islands (e.g., extending another 2 kb outward from the CpG shore).
[0110] As used herein, the term "colonic proliferative disorder" generally refers to a condition or disease involving the disordered or abnormal proliferation of colonic or rectal cells. In some instances, the disorder is selected from: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma. In some embodiments, colonic proliferative disorder includes colorectal cancer.
[0111] As used herein, the term "epigenetic parameter" generally refers to cytosine methylation. Further epigenetic parameters include, for example, histone acetylation, although they may not be directly analyzeable using the methods described, but are instead related to DNA methylation.
[0112] As used herein, the term "genetic parameter" generally refers to the mutations and polymorphisms of a gene and the sequences required for their further regulation. Examples of mutations include insertions, deletions, point mutations, inversions, and polymorphisms such as SNPs (single nucleotide polymorphisms).
[0113] As used herein, the term "hemimethylation" or "semi-methylation" generally refers to the methylation state of a palindromic CpG methylation site, in which only one of the two CpG dinucleotide sequences at the palindromic CpG methylation site is methylated (e.g., 5'-CC). M GG-3' (on-chain): 3'-GGCC-5' (off-chain)).
[0114] As used herein, the term "hypermethylation" generally refers to the average methylation state corresponding to an increase in the presence of 5-mC at one or more CpG dinucleotides in the DNA sequence of the test DNA sample, relative to the number of 5-mCs seen at the corresponding CpG dinucleotides in a normal control DNA sample. In some embodiments, the test DNA sample is derived from an individual with colonic proliferative disorder.
[0115] As used herein, the term "hypomethylation" generally refers to the average methylation state corresponding to a reduction in the number of 5-mCs seen at one or more CpG dinucleotides in the DNA sequence of the test DNA sample, relative to the number of 5-mCs seen at the corresponding CpG dinucleotides in a normal control DNA sample. In some embodiments, the test DNA sample is derived from an individual with colonic proliferative disorder.
[0116] As used herein, the term "methylation state" or "methylation condition" generally refers to the presence or absence of 5-methylcytosine ("5-mC") at one or more CpG dinucleotides in a DNA sequence. The methylation state of one or more specific CpG palindromic methylation sites (each site has two CpG dinucleotide sequences) in a DNA sequence includes "unmethylated", "fully methylated", and "hemimethylated".
[0117] As used herein, the term "methylated cytosine" generally refers to any methylated form of the cytosine base in a nucleic acid, containing a methyl or hydroxymethyl functional group at the 5' position. Methylated cytosine is known to be a regulator of gene transcription in genomic DNA. This term may include both 5-methylcytosine and 5-hydroxymethylcytosine.
[0118] As used herein, the term "methylation assay" generally refers to any assay used to determine the methylation status of one or more CpG dinucleotide sequences within a DNA sequence.
[0119] As used in this article, the term "minimal residual disease" or "MRD" generally refers to a small number of cancer cells remaining in the body after cancer treatment. MRD testing can be performed to determine the effectiveness of cancer treatment and guide further treatment planning.
[0120] As used herein, the term “MSP” (methylation-specific polymerase chain reaction (PCR)) generally refers to methylation assays, such as those described by Herman Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996 and U.S. Patent No. 5,786,146 (the contents of which are incorporated herein by reference).
[0121] As used herein, the terms "methylated transformed" or "transformed" nucleic acids generally refer to nucleic acids, such as DNA, that have undergone a DNA transformation process for methylation sequencing. Examples of transformation processes include reagent-based transformations (such as bisulfite), enzymatic transformations, or combined transformations (such as TET-assisted pyridineborane sequencing (TAPS) transformations), in which unmethylated cytosine is converted to uracil prior to PCR amplification or sequencing. Transformation processes can be used in methylation sequencing methods to distinguish between methylated and unmethylated cytosine bases.
[0122] As used in this article, the term "methylated region in cancer" generally refers to a segment of the genome containing methylation sites (CpG dinucleotides), the methylation of which is associated with malignant cell states. Regional methylation can be associated with more than one different type of cancer, or it can be specific to a single cancer type. Furthermore, regional methylation can be associated with more than one cancer subtype, or it can be specific to a single cancer subtype.
[0123] The terms "type" and "subtype" of cancer are generally used relative in this text, thus a "type" of cancer, such as breast cancer, can be a "subtype" based on, for example, stage, morphology, histology, gene expression, receptor profile, mutation profile, invasiveness, prognosis, malignancy characteristics, etc. Similarly, "type" and "subtype" can be applied at a finer level, for example, classifying a histological "type" into "subtypes," for example, based on mutation profile or gene expression. Cancer "stage" is also used to refer to cancer type classifications based on histological and pathological features associated with disease progression.
[0124] II. Analytical Samples
[0125] Cell-free biological samples can be obtained from or derived from human subjects. Cell-free biological samples can be stored under different storage conditions before processing, such as different temperatures (e.g., room temperature, refrigeration or freezing, 25°C, 4°C, -18°C, -20°C, or -80°C) or different suspensions (e.g., EDTA collection tubes, cell-free RNA collection tubes, or cell-free DNA collection tubes).
[0126] Cell-free biological samples can be obtained from subjects who have cancer, are suspected of having cancer, or do not have or are not suspected of having cancer.
[0127] Cell-free biological samples can be collected before and / or after treatment in cancer subjects. Cell-free biological samples can be obtained from subjects during treatment or a treatment regimen. Multiple cell-free biological samples can be obtained from subjects to monitor treatment efficacy over time. Cell-free biological samples can be taken from subjects who are known or suspected of having cancer but cannot obtain a definitive positive or negative diagnosis through clinical trials. Samples can be taken from subjects suspected of having cancer. Cell-free biological samples can be taken from subjects exhibiting unexplained symptoms such as fatigue, nausea, weight loss, pain, weakness, or bleeding. Cell-free biological samples can be taken from subjects with explained symptoms. Cell-free biological samples can be taken from subjects at risk of developing cancer due to factors such as family history, age, high blood pressure or prehypertension, diabetes or prediabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, alcohol consumption), or the presence of other risk factors.
[0128] Cell-free biological samples may contain one or more analytes, such as cell-free ribonucleic acid (cfRNA) molecules suitable for analysis to generate transcriptome data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for analysis to generate genomic data, or mixtures or combinations thereof. One or more such analytes (e.g., cfRNA and / or cfDNA molecules) may be isolated or extracted from one or more cell-free biological samples of the subject for downstream analysis using one or more suitable assays.
[0129] After obtaining a cell-free biological sample from the subject, the cell-free biological sample can be processed to generate a dataset indicating the subject's cancer. For example, the presence, absence, or quantification of nucleic acid molecules in the cell-free biological sample can be assessed on a locus panel of a cancer-related genome (e.g., quantitative measurement of RNA transcripts or DNA at a cancer-related genome locus). In some embodiments, processing of the cell-free biological sample obtained from the subject may include: (i) placing the cell-free biological sample under conditions sufficient to isolate, enrich, or extract multiple nucleic acid molecules; and (ii) analyzing the multiple nucleic acid molecules to generate a dataset.
[0130] In some implementations, multiple nucleic acid molecules are extracted from cell-free biological samples and sequenced to generate multiple sequencing reads. The nucleic acid molecules may include ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). Nucleic acid molecules (e.g., RNA or DNA) can be extracted from cell-free biological samples using a variety of methods, such as those derived from MP. of Reagent kit protocol, from of DNA cell-free biological mini kit, or from Norgen This is a cell-free biological DNA isolation kit protocol. The extraction method can extract all RNA or DNA molecules from a sample. Alternatively, the extraction method can selectively extract a subset of RNA or DNA molecules from a sample. RNA molecules extracted from the sample can be converted into DNA molecules via reverse transcription (RT).
[0131] Sequencing can be performed using any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), ligation sequencing, hybridization sequencing, and...
[0132] Sequencing can include nucleic acid amplification (e.g., RNA or DNA molecules). In some implementations, nucleic acid amplification is polymerase chain reaction (PCR). An appropriate number of rounds of PCR (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to adequately amplify the initial amount of nucleic acid (e.g., RNA or DNA) to the desired input level for subsequent sequencing. In some cases, PCR can be used for the whole-molecule amplification of the target nucleic acid. This can include using an aptamer sequence that can be first ligated to a different molecule, followed by PCR amplification using universal primers. PCR can be performed using any of many commercially available kits, such as those from Life Sciences. The kits provided are as follows. In other cases, only certain target nucleic acids within the nucleic acid population can be amplified. Specific primers (which may bind to aptamers) can be used to selectively amplify certain targets for downstream sequencing. PCR can include targeted amplification of one or more genomic loci, such as cancer-related genomic loci. Sequencing can include the use of simultaneous reverse transcription (RT) and polymerase chain reaction (PCR), such as those produced by... Thermo Fisher or The OneStep RT-PCR kit protocol is provided.
[0133] RNA or DNA molecules isolated or extracted from cell-free biological samples can be tagged, for example, using identifiable tags to allow multiplexing of multiple samples. Any number of RNA or DNA samples can be multiplexed. For example, a multiplexing reaction may contain RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more than 100 initial cell-free biological samples. For example, multiple cell-free biological samples can be tagged with sample barcodes so that each DNA molecule can be traced back to the sample (and object) from which the DNA molecule originated. Such tags can be ligated to RNA or DNA molecules via ligation or primer PCR amplification.
[0134] After sequencing nucleic acid molecules, the sequence reads can be subjected to appropriate bioinformatics processing to generate data indicating the presence, absence, or relative assessment of cancer. For example, sequence reads can be aligned to one or more reference genomes (e.g., genomes of one or more species, such as the human genome, e.g., hg19). The aligned sequence reads can be quantified at one or more genomic loci to generate a dataset indicating cancer. For example, quantifying sequences corresponding to multiple cancer-related genomic loci can generate a dataset indicating cancer.
[0135] Cell-free biological samples can be processed without any nucleic acid extraction. For example, cancer in a subject can be identified or monitored using probes configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes can be nucleic acid primers. The probes can have sequence complementarity with one or more nucleic acid sequences from multiple cancer-associated genomic loci or genomic regions. Multiple cancer-associated genomic loci or genomic regions may contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100 or more different cancer-associated genomic loci or genomic regions. Multiple cancer-associated genomic loci or genomic regions may contain one or more members selected from the groups listed in Tables 1-11 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, approximately 25, approximately 30, approximately 35, approximately 40, approximately 45, approximately 50, approximately 55, approximately 60, approximately 65, approximately 70, approximately 75, approximately 80, or more). Cancer-associated genomic loci or genomic regions may be associated with different cancer (e.g., colorectal cancer) stages or subtypes.
[0136] Probes can be nucleic acid molecules (e.g., RNA or DNA) that are sequence-complementary to nucleic acid sequences (e.g., RNA or DNA) of one or more genomic loci (e.g., cancer-associated genomic loci). These nucleic acid molecules can be primers or enriched sequences. Analysis of cell-free biological samples using probes selective for one or more genomic loci (e.g., cancer-associated genomic loci) can include the use of array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some implementations, DNA or RNA can be analyzed by one or more of the following: DNA / RNA isothermal amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling amplification (RCA), recombinase polymerase amplification (RPA)), immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based assays (e.g., CRISPR genotyping PCR (ctPCR), specific high-sensitivity enzyme reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR trans-reporter gene (DETECTR), and CRISPR-mediated simulated multiple event recording device (CAMERA)), and laser transmission spectroscopy (LTS).
[0137] Assay readouts can be quantified at one or more genomic loci (e.g., cancer-associated genomic loci) to generate data indicative of cancer. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple genomic loci (e.g., cancer-associated genomic loci) can generate cancer-indicating data. Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof. The assay can be a home user test configured for performance in a home environment.
[0138] In some implementations, multiplex assays can be used to simultaneously process cell-free biological samples of the subject. For example, a first assay can be used to process a first cell-free biological sample obtained or derived from the subject to generate a first dataset indicative of cancer; and a second assay, different from the first assay, can be used to process a second cell-free biological sample obtained or derived from the subject to generate a second dataset indicative of cancer. Either or both of the first and second datasets can then be analyzed to assess the subject's cancer. For example, a single diagnostic indicator or diagnostic score can be generated based on a combination of the first and second datasets. Another example is that separate diagnostic indicators or diagnostic scores can be generated based on the first and second datasets.
[0139] Cell-free biological samples can be processed using methylation-specific assays. For example, methylation-specific assays can be used to identify quantitative measures (e.g., indicating presence, absence, or relative quantity) of methylation at each of multiple cancer-related genomic loci in a subject's cell-free biological sample. Methylation-specific assays can be configured to process cell-free biological samples, such as blood or urine samples (or derivatives thereof) of a subject. Quantitative measures (e.g., indicating presence, absence, or relative quantity) of methylation at cancer-related genomic loci in a cell-free biological sample can indicate one or more cancers. Methylation-specific assays can be used to generate datasets indicating quantitative measures (e.g., indicating presence, absence, or relative quantity) of methylation at each of multiple cancer-related genomic loci in a subject's cell-free biological sample.
[0140] For example, methylation-specific assays may include one or more of the following: methylation-sensing sequencing (e.g., using bisulfite treatment), pyrosequencing, methylation-sensitive single-strand conformation analysis (MS-SSCA), high-resolution melting analysis (HRM), methylation-sensitive single nucleotide primer extension (MS-SnuPE), base-specific cleavage / MALDI-TOF, microarray-based methylation assays, methylation-specific PCR, targeted bisulfite sequencing, oxidized bisulfite sequencing, mass spectrometry-based bisulfite sequencing, or degenerate representative bisulfite sequencing (RRBS).
[0141] III. Signature Panel
[0142] This disclosure provides methods and systems for analyzing biological samples to obtain measurable features from combinations of highly methylated regions in DNA associated with the development of colonic proliferative disease, thereby identifying a signature panel of regions. Features from the signature panel can be processed using a trained algorithm (e.g., a machine learning model) to create a classifier configured to stratify individual populations of colonic proliferative disease. The method is characterized by using one or more nucleic acids having methylated regions described in the signature panel, which are contacted prior to sequencing with one or more reagents capable of distinguishing between methylated and unmethylated CpG dinucleotides within the identified regions.
[0143] The signature panels described herein generally refer to a collection of genomic DNA target regions identified in cell-free nucleic acid samples that exhibit increased cytosine base methylation in samples associated with colonic proliferative disease. The formation of signature panels allows for rapid and specific analysis of specific methylated regions associated with colonic proliferative disease. The signature panels described and used in the methods described herein can be used to improve the diagnosis, prognosis, treatment selection, and monitoring (e.g., treatment surveillance) of colonic proliferative disease.
[0144] The signature panel and method disclosed herein offer significant improvements over current methods in addressing the need for biomarkers or signature panels used to detect early colorectal proliferative disorders from bodily fluid samples such as whole blood, plasma, or serum. Current methods for detecting and diagnosing colorectal proliferative disorders include colonoscopy, sigmoidoscopy, and fecal occult blood tests for colorectal cancer. Compared to these methods, the method provided herein is significantly less invasive than colonoscopy and is at least as sensitive as or more sensitive than sigmoidoscopy, fecal immunochemical testing (FIT), and fecal occult blood testing (FOBT). The method provided herein offers significant advantages in sensitivity and specificity compared to these currently used biomarkers due to the advantageous combination of gene panels with highly sensitive assay techniques.
[0145] In some embodiments, the methylated region in the cancer includes CpG islands. In some embodiments, the methylated region in the cancer includes CpG shores. In some embodiments, the methylated region in the cancer includes CpG racks. In some embodiments, the methylated region in the cancer includes both CpG islands and CpG shores. In some embodiments, the methylated region in the cancer includes CpG islands, CpG shores, and CpG racks.
[0146] In some implementations, the methylated region in cancer includes CpG islands and approximately 0 to 4 kilobase (kb) of upstream and downstream sequences. The methylated region in cancer may also include CpG islands and the following sequences: approximately 0 to 3 kb upstream and downstream, approximately 0 to 2 kb upstream and downstream, approximately 0 to 1 kb upstream and downstream, approximately 0 to 500 base pairs (bp) upstream and downstream, approximately 0 to 400 bp upstream and downstream, approximately 0 to 300 bp upstream and downstream, approximately 0 to 200 bp upstream and downstream, or approximately 0 to 100 bp upstream and downstream.
[0147] Based on several examples, many design parameters can be considered when selecting hypermethylated regions in cancer. In some instances, the length of the methylated regions is approximately 200 bp, 300 bp, 400 bp, or 500 bp. Data for this selection process can be obtained from various sources, such as The Cancer Genome Atlas (TCGA) (cancergenome.nih.gov), using, for example, data used for a wide range of cancers. The methylation value is derived from the Infinium HumanMethylation450 BeadChip or obtained from other sources based on bisulfite whole-genome sequencing or other methods. In some embodiments, regions can be selected using "methylation values" (which can be derived from TCGA Level 3 methylation data, which in turn are derived from beta values of approximately -0.5 to 0.5). In some embodiments, amplification is performed using a primer set designed to amplify at least one methylation site with a methylation value normally below approximately -0.3. This can be established in multiple normal tissue samples, such as approximately 4. Methylation values can be equal to or below approximately -0.1, approximately -0.2, approximately -0.3, approximately -0.4, approximately -0.5, approximately -0.6, approximately -0.7, approximately -0.8, approximately -0.9, or approximately -1.0.
[0148] In some embodiments, the primer set is designed to amplify at least one methylation site whose average methylation value differs from that of normal tissue by more than a predefined threshold, such as about 0.3. In some embodiments, this difference may be greater than about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0. In some instances, the proximity of other methylation sites satisfying this requirement may also play a role in the selected region. In some embodiments, the primer set includes primer pairs that amplify at least one methylation site within about 200 bp, whose methylation value is also below about -0.3 under normal conditions, and whose average methylation value differs from that of normal tissue by more than about 0.3.
[0149] In some instances, a target region is selected if the methylation of a region is greater than that of the same region in samples obtained from or derived from one or more healthy individuals (e.g., individuals without cancer). This selection can be performed manually or computationally. In some instances, a region is selected if it has at least about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 100%, or more than about 100% more methylation than a sample from a healthy individual. In another instance, a region can be selected if the number of reads in a disease sample mapped to a region with a predefined threshold methylation CpG count exceeds the same predefined threshold methylation CpG count in the same region of a healthy individual sample. For a given region, the methylated CpG count used as a baseline threshold in a healthy sample can vary, but a read that maps to that region exceeding the baseline threshold of the methylated CpG count for that region in a healthy sample can indicate an important region, regardless of how the threshold CpG count fluctuates.
[0150] In some instances, target regions can be selected for amplification based on the number of samples with methylation at that site, as observed in validation studies. For example, a region can be selected if samples tested from diseased individuals have at least approximately 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% methylation compared to samples from healthy individuals. For instance, regions can be selected if they are methylated in at least approximately 75% of the tested tumors (including within a specific subtype). For some validations, tumor-derived cell lines can be used for testing.
[0151] This disclosure also provides a method for performing an assay to determine genetic and / or epigenetic parameters of one or more genes selected from the signature panel described herein, as well as their promoters and regulatory elements. In some embodiments, the assay, performed according to the method, is for detecting methylation within one or more genes selected from the signature panel described herein, wherein the methylated nucleic acid is present in a solution also containing an excess of background DNA, wherein the background DNA is present at a concentration of about 100 to 1000, about 100 to 10000, about 100 to 10000, about 1000 to 100000, or about 10000 to 100000 times the concentration of the DNA to be detected. In some embodiments, the concentration of the DNA to be detected is greater than about 100,000 times the concentration of the background DNA. In some embodiments, the method includes contacting a nucleic acid sample obtained from a subject with at least one reagent or a series of reagents (e.g., a reagent for distinguishing between methylated and unmethylated CpG dinucleotides within the target nucleic acid).
[0152] The tumors or proliferative colorectal diseases described herein may be selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some embodiments, proliferative colorectal diseases include colorectal cancer.
[0153] A signature panel containing informative methylation regions can be selected based on the intended assay purpose. For targeting methods, primer pairs can be designed based on the intended set of target regions. In some embodiments, the region set includes at least one, at least two, at least three, or more than three regions listed in Table 1. In some embodiments, the region set includes all regions listed in Table 1.
[0154] In some implementations, the set of methyl regions associated with colorectal cancer is selected from Table 1.
[0155] In some embodiments, the cancer panel includes regions selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, SFMBT2, CHST10, CCNA1, BEND4, KRBA1, S1PR1, PPP1R16B, IKZF1, LONRF2, ZFP82, and FLT3 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes all regions listed in Table 1. In some embodiments, the probe points to sequences selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, ZNF543, SFMBT2, CHST10, CCNA1, BEND4, KRBA1, S1PR1, PPP1R16B, IKZF1, LONRF2, ZFP82, and FLT3.
[0156] Table 1
[0157]
[0158]
[0159] In some embodiments, the method further includes quantifying methylation signals, wherein values exceeding a predetermined threshold indicate colonic proliferative disease. In some embodiments, the quantification and comparison of each methylation site in colonic proliferative disease are performed independently. Therefore, a count of tumor-positive signals can be established for each site. In some embodiments, the method further includes determining the proportion of sequencing reads containing tumor signals, wherein a proportion exceeding a threshold indicates colonic proliferative disease. In some embodiments, the determination of each methylation site in colonic proliferative disease is performed independently.
[0160] As used herein, the term "threshold" generally refers to a value selected to identify, separate, or distinguish between two groups of objects. In some implementations, the threshold distinguishes methylation status into disease (e.g., malignancy) and non-disease (e.g., healthy) states. In some implementations, the threshold may distinguish different stages of a disease (e.g., stage 1, stage 2, stage 3, or stage 4). The threshold may be set based on the relevant disease and may be determined based on earlier analyses, such as analysis of a training set, or calculated based on a set of inputs with known characteristics (e.g., healthy, disease, or disease stage). Thresholds may also be set for gene regions based on the predicted methylation value at a specific site. The threshold may be different for each methylation site, and data from multiple sites may be combined in the final analysis.
[0161] In some embodiments of the above method, the cancer panel comprises at least one, at least two, at least three, or more than three regions selected from the following: ITGA4, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 2. In some embodiments, the probe points to at least one, at least two, at least three, or more than three sequences selected from the following: ITGA4, TMEM163, SFMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.
[0162] Table 2
[0163]
[0164]
[0165] In some embodiments, the cancer panel includes regions selected from at least one, at least two, at least three, or more than three of the following: EMBP1, TMEM163, SMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes one or more regions listed in Table 3. In some embodiments, the probe points to sequences selected from at least one, at least two, at least three, or more than three of the following: EMBP1, TMEM163, SMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, S1PR1, and PPP1R16B.
[0166] Table 3
[0167]
[0168]
[0169] In some embodiments, the cancer panel includes regions selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, and S1PR1, and the tumor is colorectal cancer. In some embodiments, the cancer panel includes one or more regions listed in Table 4. In some embodiments, the probe targets sequences selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SMBT2, ELMO1, ZNF543, CHST10, CCNA1, BEND4, KRBA1, and S1PR1.
[0170] Table 4
[0171]
[0172]
[0173] In some embodiments, the cancer panel includes regions selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, and ZNF543, and the tumor is colorectal cancer. In some embodiments, the cancer panel includes regions listed in Table 5. In some embodiments, the probe targets sequences selected from at least one, at least two, at least three, or more than three of the following: ITGA4, EMBP1, TMEM163, SFMBT2, ELMO1, and ZNF5431.
[0174] Table 5
[0175] Methyl region (gene ID; chromosome: start-end position) ITGA4;chr2:181457004-181457950 EMBP1;chr1:121519076-121519744 TMEM163;chr2:134718243-134719428 SFMBT2;chr10:7408046-7408953 ELMO1;chr7:37448612-37449471 ZNF543;chr19:57320164-57320845
[0176] In some embodiments, the cancer panel includes one or more of regions ITGA4 and EMBP1 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes one or more regions listed in Table 6. In some embodiments, the probe is directed to sequences including ITGA4 and EMBP1.
[0177] Table 6
[0178] Methyl region (gene ID; chromosome: start-end position) ITGA4;chr2:181457004-181457950 EMBP1;chr1:121519076-121519744
[0179] In some embodiments of the above method, the cancer panel includes at least one, at least two, at least three, or more regions selected from the following: KZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, IRF4, LOC100130992, CHST11, CHST15, RASSF2, EMILIN2, TMEM163, CHST10, and HCK (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes one or more regions listed in Table 7. In some implementations, the probe points to a sequence selected from at least one, at least two, at least three, or more than three of the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, IRF4, LOC100130992, CHST11, CHST15, RASSF2, EMILIN2, TMEM163, CHST10, and HCK.
[0180] Table 7
[0181]
[0182]
[0183] In some embodiments of the above method, the cancer panel includes at least one, at least two, at least three, or more regions selected from the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, and IRF4 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes one or more regions listed in Table 8. In some implementations, the probe points to a sequence selected from at least one, at least two, at least three, or more than three of the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, ST3GAL1, ZEB2 NR3C1, ITGA4, GALNT14, CHST11, PPP1R16B, MGAT3, ZNF264, BEND4, and IRF4.
[0184] Table 8
[0185]
[0186]
[0187] In some embodiments of the above method, the cancer panel comprises at least one, at least two, at least three, or more regions selected from the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 9. In some embodiments, the probe points to a sequence selected from at least one, at least two, at least three, or more than three of the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1.
[0188] Table 9
[0189]
[0190]
[0191] In some embodiments of the above methods, the cancer panel comprises regions selected from at least one, at least two, at least three, or more than three of the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel comprises one or more regions listed in Table 10. In some embodiments, the probe targets sequences selected from at least one, at least two, at least three, or more than three of the following: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1.
[0192] Table 10
[0193] Methyl region (gene ID; chromosome: start-end position) IKZF1;chr7:50303445-50305526 KCNQ5;chr6:72620772-72623556 ELMO1;chr7:37447220-37450201 CHST2;chr3:143118680-143121423 PRKCB;chr16:23835445-23837405 FLI1;chr11:128691887-128696541
[0194] In some embodiments of the above methods, the cancer panel includes regions selected from at least one, at least two, or at least three of the following: IKZF1, KCNQ5, and ELMO1 (e.g., where the tumor is colorectal cancer). In some embodiments, the cancer panel includes one or more regions listed in Table 11. In some embodiments, the probe targets sequences selected from at least one, at least two, or at least three of the following: IKZF1, KCNQ5, and ELMO1.
[0195] Table 11
[0196] Methyl region (gene ID; chromosome: start-end position) IKZF1;chr7:50303445-50305526 KCNQ5;chr6:72620772-72623556 ELMO1;chr7:37447220-37450201
[0197] On one hand, this disclosure provides a method for identifying methylation signatures indicating biological characteristics, the method comprising: acquiring data for a population containing multiple genomic methylation datasets associated with a colonic proliferative disease state, each of the genomic methylation datasets being associated with biological information of a corresponding sample; separating the methylation datasets into a first group corresponding to a tissue or cell type having the biological characteristics and a second group corresponding to multiple tissue or cell types not having the biological characteristics; matching the methylation data of the first group with the methylation data of the second group at site-by-site in the genome; identifying a set of CpG sites at site-by-site in the genome, these sites satisfying a predetermined threshold for establishing differential methylation between the first group and the second group; using the CpG site set to identify target genomic regions containing at least one, at least two, at least three, or more than three differentially methylated CpGs within approximately 30 to 300 bp that satisfy the predetermined criteria, to identify differentially methylated genomic regions thereby providing methylation signatures indicating biological characteristics associated with the presence of colonic proliferative disease.
[0198] In some instances, the target genomic region contains at least one, at least two, at least three, or more than three differentially methylated CpG sites within regions of the following lengths: approximately 30 to 150 bp, approximately 40 to 150 bp, approximately 50 to 150 bp, approximately 75 to 150 bp, approximately 100 to 150 bp, approximately 150 to 300 bp, approximately 150 to 250 bp, approximately 150 to 200 bp, approximately 200 to 300 bp, or approximately 250 to 300 bp.
[0199] In some instances, the target genomic region contains at least four differentially methylated CpG sites, at least four differentially methylated CpG sites, at least five differentially methylated CpG sites, at least six differentially methylated CpG sites, at least seven differentially methylated CpG sites, at least eight differentially methylated CpG sites, at least nine differentially methylated CpG sites, at least ten differentially methylated CpG sites, at least 12 differentially methylated CpG sites, or at least 15 differentially methylated CpG sites.
[0200] In some embodiments, the method further includes verifying the extended target genome region by detecting differential methylation within the extended target genome region using DNA from at least one independent sample possessing the biological trait and DNA from at least one independent sample not possessing the biological sample.
[0201] In some embodiments, the identification further includes limiting the CpG site set to CpG sites that further exhibit differential methylation compared to peripheral blood mononuclear cells from reference or control samples.
[0202] In some implementations, a predetermined threshold of at least about 50% methylation is set in the first group.
[0203] In some implementations, the predetermined threshold is the average methylation difference between the first group and the second group of at least about 0.3.
[0204] In some implementations, biological characteristics include malignancy.
[0205] In some implementations, biological traits include cancer type.
[0206] In some implementations, biological traits include cancer stage.
[0207] In some implementations, biological traits include cancer classification.
[0208] In some implementation schemes, cancer classification includes cancer grading.
[0209] In some implementation schemes, cancer classification includes histological classification.
[0210] In some implementations, biological traits include metabolic profiles.
[0211] In some implementations, biological traits include mutations.
[0212] In some implementations, the mutation is a disease-related mutation.
[0213] In some implementations, biological characteristics include clinical outcomes.
[0214] In some implementations, biological traits include drug response.
[0215] In some embodiments, the method further includes designing multiple PCR primer pairs to amplify portions of an extended target genomic region, each portion containing at least one differentially methylated CpG site.
[0216] In some implementations, the design of multiple primer pairs includes converting unmethylated cytosine to uracil to mimic the conversion of cytosine to uracil, and designing primer pairs using the converted sequences.
[0217] In some implementations, primer pairs are designed to be methylation-prone.
[0218] In some implementations, the primer pairs are methylation-specific.
[0219] In some implementations, the primer pairs contain no CpG residues and have no preference for methylation state.
[0220] On the one hand, this disclosure provides a method for synthesizing primer pairs that are specific to methylation signatures, the method comprising: performing the method of this disclosure and synthesizing the designed primer pairs.
[0221] IV. Nucleic acid transformation and methylation sequencing
[0222] A. Nucleic acid processing
[0223] Methylation sequencing can utilize a variety of methods, including chemical and enzymatic conversions of nucleic acid bases, to distinguish methylated cytosine from unmethylated cytosine in a nucleic acid sequence. These assays allow determination of the methylation status of one or more CpG dinucleotides (e.g., CpG islands) within a DNA sequence. Such assays may include, in particular, DNA sequencing of bisulfite-treated or enzyme-treated DNA, polymerase chain reaction (PCR) (for sequence-specific amplification), quantitative PCR (qPCR), or digital droplet PCR (ddPCR), and DNA blot analysis. In various instances, DNA in a biological sample is treated in such a way that the unmethylated cytosine base at the 5' position is converted to uracil, thymine, or another base that differs from cytosine in hybridization behavior. This can be termed a "conversion."
[0224] In some implementations, the reagent converts the unmethylated cytosine base at the 5'-position into uracil, thymine, or another base that differs from cytosine in hybridization behavior.
[0225] DNA bisulfite modification generally refers to tools used to assess CpG methylation status. A common method for analyzing the presence of 5-methylcytosine (5-mC) in DNA is based on the reaction of bisulfite with cytosine, followed by alkaline desulfurization, which converts cytosine to uracil, corresponding to the base-pairing behavior of thymine. For example, genome sequencing has been adapted to analyze DNA methylation patterns and 5-methylcytosine distribution using bisulfite treatment (e.g., as described by Frommer et al., Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992, the contents of which are incorporated herein by reference). However, it is noteworthy that 5-methylcytosine remains unmodified under these conditions. Therefore, the original DNA is transformed in such a way that methylcytosine (methyl-C), initially indistinguishable from cytosine by hybridization behavior, can now be detected as the only remaining cytosine using various molecular biology techniques, such as amplification and hybridization or sequencing. In each instance, other reagents can achieve the same results as bisulfite modification suitable for methylation sequencing.
[0226] A commonly used direct sequencing method uses bisulfite-treated DNA that has been amplified by PCR, and it is suitable for whole-genome bisulfite sequencing (WGBS) or targeted bisulfite sequencing.
[0227] Targeted bisulfite sequencing refers to commercially available NGS methods used to assess site-specific changes in DNA methylation. Probes are designed to be both strand-specific and bisulfite-specific. Both methylated and unmethylated sequences are amplified. This process is similar to pyrosequencing but generally offers higher throughput. In some implementations, next-generation sequencing platforms are used to deliver large amounts of useful DNA methylation information (e.g., EPIGENTEK, Farmingdale, NY, and ZYMORESEARCH, Irvine, CA). By bisulfite-treating DNA, followed by PCR amplification of the target region, library construction, and sequencing of the amplicon regions, single-base resolution methylation analysis of individual cytosines in DNA can be facilitated. Specific primers can be designed for the target region, and changes in cytosine methylation within that region can be assessed. Each target DNA methylation site can be evaluated at high sequencing coverage depth to obtain accurate, quantitative, and single-base resolution data output.
[0228] Enzymatic methylation sequencing (EM-seq) relies on the enzymatic transformation of nucleic acids for methylome analysis. Data suggests that the process of generating EM-seq libraries does not damage DNA as much as bisulfite sequencing. EM-seq libraries achieve higher PCR yields despite using fewer PCR cycles for all DNA inputs, indicating less DNA loss during enzymatic processing and library preparation compared to whole-genome bisulfite sequencing (WGBS). Conversely, fewer PCR cycles translate to more complex libraries and fewer PCR replicas during sequencing. The average insert size of EM-seq libraries can also be larger than that of WGBS, further supporting the fact that DNA remains intact. In the EM-seq process, TET2 oxidizes 5-mC and 5-hmC, preventing APOBEC deamination in the next step. Instead, unmodified cytosine is deamination to uracil. In some embodiments, the targeting method includes enzymatic transformation of nucleic acids (TEM-seq). In some embodiments, the methylation sequencing method uses… The results were obtained using Enzymatic Methyl-seq (New England Biolabs, Ipswich, MA), which is useful for the identification of 5mC and 5hmC.
[0229] In another instance, 5hmC can also be sequenced using TET-assisted bisulfite sequencing (TAB-seq) (e.g., as described by Yu, M. et al. (2012). Nat. Protoc. 7, 2159-2170, the contents of which are incorporated herein by reference) (WiseGene; The DNA fragments were detected using a sequence of T4 phage β-glucosyltransferases (T4-BGT), followed by treatment with a 10⁻¹¹ translocation (TET) dioxygenase prior to the addition of sodium bisulfite. T4-BGT glycosylated 5hmC to form β-glucosyl-5-hydroxymethylcytosine (5ghmC), which was then oxidized to 5caC by TET. Only 5ghmC was unaffected by the subsequent deamination by sodium bisulfite, allowing 5hmC to be distinguished from 5mC by sequencing.
[0230] Oxidized bisulfite sequencing (oxBS) provides another method for distinguishing 5mC from 5hmC (e.g., as described by Booth, MJ, et al., 2012 Science 336:934-937, the contents of which are incorporated herein by reference). The oxidizing agent potassium perruthenate converts 5hmC to 5-formylcytosine (5fC), and subsequent sodium bisulfite treatment deamination of 5fC to generate uracil. 5mC remains unchanged, and therefore this method can be used for identification.
[0231] APOBEC-coupled epigenetic sequencing (ACE-seq) completely excludes bisulfite conversion and relies on enzymatic conversion to detect 5hmC (e.g., as described by Schutsky, EK et al., Nat. Biotechnol., 2018 Oct 8, the contents of which are incorporated herein by reference). In this method, T4-BGT glycosylates 5hmC to 5ghmC and protects it from deamination by apolipoprotein B mRNA editing enzyme subunit 3A (APOBEC3A). Cytosine and 5mC are deaminated by APOBEC3A and sequenced to thymine.
[0232] In another example, a bisulfite-free and base-level resolution sequencing method, namely TET-assisted pyridineborane sequencing (TAPS), can be used for the detection of 5mC and 5hmC. TAPS combines the 10-11 translocation (TET) oxidation of 5mC and 5hmC to 5-carboxycytosine (5caC) with the pyridineborane reduction of 5caC to dihydrouracil (DHU). Subsequent PCR converts DHU to thymine, achieving the C-to-T conversion of 5mC and 5hmC. TAPS directly detects the modification with high sensitivity and specificity without affecting the unmodified cytosine. (e.g., as described by Liu, Y. et al. Nat Biotechnol. 2019 Apr; 37(4):424-429, the contents of which are incorporated herein by reference).
[0233] TET-assisted 5-methylcytosine sequencing (TAmC-seq) enriched the 5mC locus using two consecutive enzymatic reactions followed by affinity pull-down (e.g., as described by Zhang, L. 2013, Nat Commun 4:1517, the contents of which are incorporated herein by reference). Fragment DNA was treated with T4-BGT to protect 5hmC via glycosylation. 5mC was then oxidized to 5hmC using the mTET1 enzyme, and the newly formed 5hmC was tagged with T4-BGT using a modified glucose moiety (6-N3-glucose). Click chemistry was used to introduce the biotin tag, enabling the enrichment of DNA fragments containing 5mC for detection and genome-wide profiling.
[0234] B. Next-generation sequencing
[0235] In some implementations, sequencing reads are generated via next-generation sequencing. This allows for higher read depths for a given region. These can be high-throughput methods, including, for example... (Solexa) sequencing, DNB-Sequencer T7 Or G400 (MGI Tech Co., Ltd), Sequencing (GenapSys, Inc.), Roche 454 sequencing (Roche Sequencing Solutions, Inc.), Ion Torrent sequencing (Thermo Fisher Scientific), and SOLiD sequencing (Thermo Fisher Scientific). The number of sequencing reads can be adjusted based on the amount of DNA input and the depth of data required for analysis.
[0236] In some implementations, sequencing reads are generated simultaneously from samples obtained from multiple patients, with each patient's cell-free nucleic acid fragment labeled with a barcode. This allows for parallel analysis of multiple patients in a single sequencing run.
[0237] In another aspect, this disclosure provides a kit for detecting tumors, comprising reagents for performing the methods described above and instructions for detecting tumor signals. The reagents may include, for example, primer sets, PCR reaction components, and / or sequencing reagents.
[0238] C. Targeted sequencing
[0239] In targeted methylation sequencing methods, the target region in a biological sample (such as cfDNA) is analyzed to determine the methylation status of the target gene sequence. In some implementations, the target region includes adjacent nucleotides of the target region (such as at least about 16 adjacent nucleotides of the target region) or hybridizes with it under stringent conditions. In various instances, targeted sequencing can be achieved using hybridization capture and amplicon sequencing methods.
[0240] D. Hybrid capture
[0241] The hybridization methods described herein can be used for various forms of nucleic acid hybridization, such as in-solution hybridization and hybridization on solid supports (e.g., RNA, DNA, and in situ hybridization on membranes, microarrays, and cell / tissue slides). Specifically, the methods are suitable for in-solution hybridization capture for target enrichment of certain types of genomic DNA sequences (e.g., exons) used in next-generation targeted sequencing. For hybridization capture methods, cell-free nucleic acid samples undergo library preparation. As used herein, “library preparation” includes end repair, A-tailing, aptamer ligation, or any other preparation of cell-free DNA to allow subsequent DNA sequencing. In some instances, the prepared cell-free nucleic acid library sequences contain aptamers, sequence tags, or index barcodes linked to the cell-free nucleic acid sample molecules. Various commercially available kits can be used to aid library preparation for next-generation sequencing methods. The construction of next-generation sequencing libraries may include the preparation of nucleic acid targets using a series of coordinated enzymatic reactions to produce a set of random DNA fragments of a specific size for high-throughput sequencing. Advances and developments in various library preparation technologies have expanded the application of next-generation sequencing in fields such as transcriptomics and epigenetics.
[0242] Improvements in sequencing technology have led to changes and improvements in library preparation. These include, for example, […].
[0243] Bioo Kapa New England Life Pacific and The company's next-generation sequencing library preparation kit provides consistency and reproducibility for a variety of molecular biology reactions, ensuring compatibility with the latest NGS instrument technologies.
[0244] In different instances of targeted gene capture panels, various library preparation kits can be selected from Nextera Flex. DNA Prep Ion (ThermoFisher ), (Thermo Fisher Agilent ClearSeq Capture Bioo xGen and
[0245] In some embodiments, a hybridization capture method is performed on the prepared library sequences using specific probes. In some embodiments, the term "specific probe," as used herein, generally refers to a probe specific to known methylation sites. In some embodiments, the specific probe is designed based on using the human genome as a reference sequence and using specific genomic regions known to have methylation sites as target sequences. Specifically, genomic regions known to have methylation sites may include at least one of the following regions: promoter regions, CpG island regions, CGI island-shore regions, and imprinted gene regions. Therefore, when hybridization capture is performed using specific probes of some embodiments, sequences in the sample genome that are complementary to the target sequence can be effectively captured, for example, regions in the sample genome known to have methylation sites (also referred to herein as "specific genomic regions").
[0246] According to one example, the methylated regions described herein are used to design specific probes. In some embodiments, specific probes are designed using commercially available methods (e.g., the eArray system). The probe length is sufficient to hybridize with the target methylated region with adequate specificity. In various examples, the probes are 10-mers, 11-mers, 12-mers, 13-mers, 14-mers, 15-mers, 16-mers, 17-mers, 18-mers, 19-mers, or 20-mers.
[0247] The regions listed in Tables 1-11 above were screened using database resources (such as gene ontologies). Based on the principle of complementary base pairing, single-stranded capture probes can combine complementaryly with single-stranded target sequences to successfully capture the target region. In some embodiments, the designed probes can be designed as solid capture chips (where the probes are fixed on a solid support) or as liquid capture chips (where the probes are free in a liquid). However, due to various limitations such as probe length, probe density, and high cost, solid capture chips are rarely used, while liquid capture chips are more commonly used.
[0248] In some implementations, GC-rich sequences (with GC base content exceeding 60%) may exhibit reduced capture efficiency due to the molecular structure of C and G bases, compared to normal sequences (where the average content of A, T, C, and G bases is 25%). For key study regions, such as CGI regions (CpG islands), it may be recommended to design a larger number of probes to obtain sufficient and accurate CGI data.
[0249] E. Amplicon-based sequencing
[0250] The transformed DNA fragment can be amplified. In some embodiments, amplification is performed using primers designed to anneal a methylated transformation target sequence having at least one methylation site. Methylation sequencing transformation results in the conversion of unmethylated cytosine to uracil, while 5-methylcytosine remains unaffected. Therefore, the "transformation target sequence" is understood to be a sequence in which cytosine known to be a methylation site is fixed as "C" (cytosine), and cytosine known to be unmethylated is fixed as "U" (uracil; in primer design, this can be considered as "T" (thymine)).
[0251] In various embodiments, the DNA source is cell-free DNA from whole blood, plasma, serum, or genomic DNA extracted from cells or tissues. In some embodiments, the length of the amplified fragment is between about 100 and 200 base pairs. In some embodiments, the DNA source is extracted from a cellular source (e.g., tissue, biopsy, cell line), and the length of the amplified fragment is between about 100 and 350 base pairs. In some embodiments, the amplified fragment contains at least one 20-base-pair sequence comprising at least one, at least two, at least three, or more than three CpG dinucleotides. Amplification can be performed using a set of primer oligonucleotides according to this disclosure and a thermostable polymerase can be used. Amplification of several DNA segments can be performed simultaneously in the same reaction vessel. In some embodiments, two or more fragments are amplified simultaneously. For example, amplification can be performed using polymerase chain reaction (PCR).
[0252] Primers designed to target these sequences may exhibit a degree of preference for the transformed methylated sequences. In some implementations, PCR primers are designed to be methylation-specific for targeted methylation sequencing applications. This allows for higher sensitivity in some applications. For example, primers may be designed to contain a distinguishable nucleotide (specific to the methylated sequence after bisulfite conversion) that is positioned to achieve optimal distinguishing (e.g., in PCR applications). The distinguishing nucleotide may be located at the 3' end or the penultimate position.
[0253] In some implementations, primers are designed to amplify DNA fragments of 75 to 350 bp in length. This is a known general size range for circulating DNA, and, according to this example, optimizing primer design to account for the target size can improve the sensitivity of the method. Primers may be designed to amplify regions of approximately 50 to 200, approximately 75 to 150, or approximately 100 or 125 bp in length.
[0254] In some embodiments of the methods described herein, methylation-specific primer oligonucleotides can be used to detect the methylation status of a preselected CpG position in a nucleic acid sequence using an amplicon-based method. Amplifying bisulfite-treated DNA using methylation-specific primers allows differentiation between methylated and unmethylated nucleic acids. MSP primer pairs contain at least one primer that hybridizes to the transformed CpG dinucleotide. Therefore, the primer sequence contains at least one CpG, TpG, or CpA dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the 3' position of the C in the CpG. Therefore, the primer base sequence may need to contain a sequence of at least 18 nucleotides in length that hybridizes to the pretreated nucleic acid sequence and its complementary sequence, wherein the oligomer base sequence contains at least one CpG, TpG, or CpA dinucleotide. In some embodiments, the MSP primers contain 2 to 5 CpG, TpG, or CpA dinucleotides. In some embodiments, the dinucleotide is located within the 3' half of the primer; for example, for a primer of 18 bases in length, the specified dinucleotide is located within the first 9 bases from the 3' end of the molecule. In addition to CpG, TpG, or CpA dinucleotides, primers may also contain several methyl conversion bases (e.g., cytosine to thymine, or guanine to adenosine on the hybrid strand). In some embodiments, primers are designed to contain no more than two cytosine or guanine bases.
[0255] In some implementations, each region is amplified using multiple primers. In some implementations, these regions do not overlap. These regions can be directly adjacent or spaced apart (e.g., spaced at intervals of 10, 20, 30, 40, or 50 bp). Since target regions (including CpG islands, CpG shores, and / or CpG skeletons) are typically longer than 75 to 150 bp, this example allows for the evaluation of methylation status at more (or all) sites across a given target region.
[0256] Primers can be designed for target regions using suitable tools such as Primer3, Primer3Plus, and Primer-BLAST. As discussed, bisulfite conversion results in the conversion of cytosine to uracil and 5'-methylcytosine to thymine. Therefore, primer localization or targeting can utilize the methylated sequences resulting from bisulfite conversion, depending on the desired level of methylation specificity.
[0257] The amplified target region is designed to have at least 10 CpG dinucleotide methylation sites. However, in some instances, amplifying regions with more than 10 CpG methylation sites may be advantageous. For example, a 300 bp long sequence read may have approximately 10, 20, 30, 40, or 50 CpG methylation sites methylated in nucleic acid samples associated with colonic proliferative disease. In various instances, the methylated regions identified in Tables 1-11 may have at least 25, 50, 100, 200, 300, 400, or 500 CpG methylation sites methylated in nucleic acid samples associated with colonic proliferative disease. In some embodiments, primers are designed to amplify DNA fragments containing 3 to 20 CpG methylation sites in the target region. Overall, this approach allows for the querying of more methylation sites in a single sequencing read and provides additional certainty (excluding false positives) because multiple consistent methylations may be detected in a single sequencing read. In some implementations, the tumor signal comprises more than two methylated regions selected from Tables 1-11. In this example, detecting multiple tumor signals can improve the confidence of tumor detection. Such signals may be at the same site or at different sites. In some implementations, the detection of more than one tumor signal in the same region indicates a tumor.
[0258] In some implementations, the number of CpG sites in identified methylated regions can be modeled between two populations with different characteristics of colonic cell proliferative disorder to identify a methylation threshold in which the number of CpG sites in one region exceeds the threshold to indicate colonic cell proliferative disorder.
[0259] In each instance, the number of CpG sites indicating colorectal cancer in the identified methylated regions is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18, wherein the presence of methylated CpGs exceeding this number indicates colorectal cancer and can be used as input features for machine learning models that serve as classifiers to stratify a population into healthy individuals and individuals with colorectal cancer.
[0260] In this example, the detection of multiple tumor signals indicating methylation at the same site in the genome can improve the confidence of tumor detection. The detection of methylation at adjacent sites in the genome, even if the signals come from different sequencing reads, can also improve the confidence of tumor detection. This reflects another type of signal consistency. In some embodiments, the detection of adjacent or overlapping tumor signals in at least two different sequencing reads indicates a tumor. In some embodiments, adjacent or overlapping tumor signals are within the same CpG island. In some embodiments, the detection of 3 to 34 proximal methylation sites in a cell-free DNA fragment indicates a tumor. In some embodiments, the detection of 3 to 34 methylated CpG sites in a fragment is used to identify a threshold to distinguish individual groups with certain characteristics (e.g., healthy, diseased, or disease stage). In some implementations, the detection of approximately 4 to 10, approximately 4 to 15, approximately 10 to 20, approximately 15 to 20, approximately 15 to 25, approximately 20 to 25, approximately 20 to 34, approximately 25 to 34, or approximately 30 to 34 methylated proximal CpG sites in a read fragment is used to determine a threshold to distinguish individual groups with certain characteristics (e.g., health, disease, or disease stage). As used herein, the term "proximal CpG site" refers to a CpG site on the same nucleic acid fragment in a cell-free nucleic acid sample that is adjacent to or between 2 to 10 CpG sites.
[0261] In some implementations, amplification is performed using more than 100 primer pairs. Amplification can be performed using approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, or more primer pairs. In some implementations, amplification is multiplex amplification. Multiplex amplification allows for the parallel collection of large amounts of methylation information from many target regions of the genome, even from cfDNA samples that are typically not DNA-rich. Multiplex amplification can be extended to a platform such as Ion... Up to approximately 24,000 amplicons can be queried simultaneously. In some implementations, amplification is nested amplification. Nested amplification improves sensitivity and specificity.
[0262] In addition, another rapid and robust approach for examining multiple methylated sequences in parallel is known as simultaneous targeted methylation sequencing (sTM-Seq). Key features of this technology include the elimination of the need for large amounts of high-molecular-weight DNA and nucleotide-specific differentiation between 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC). Furthermore, sTM-Seq is scalable and can be used to investigate multiple loci in dozens of samples in a single sequencing run. Freely available web-based software and universal primers for multipurpose barcoding, library preparation, and custom sequencing make sTM-Seq affordable, efficient, and versatile (e.g., as described by Asmus, N. et al., Curr Protoc Hum Genet. 2019 Apr; 101(1), the contents of which are incorporated herein by reference).
[0263] Generally, the methods and systems presented herein are useful for preparing cell-free polynucleotide sequences for downstream sequencing applications. In some implementations, the sequencing method is classic Sanger sequencing. Sequencing methods may include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, and RNA-Seq. DigitalGene Expression Next-generation sequencing, single-molecule synthesis sequencing (SMSS) Massive parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.
[0264] Pyrosequencing refers to a real-time sequencing technology based on the photometric detection of pyrosequencing released after nucleotide incorporation. It is suitable for simultaneously analyzing and quantifying the methylation levels at several CpG sites. After genomic DNA transformation, the target region is amplified using polymerase chain reaction (PCR), in which one of the two primers is biotinylated. The PCR-generated template is single-stranded, and the pyrosequencing primers are annealed to quantitatively analyze CpG sites. After bisulfite treatment and PCR, the degree of methylation at each CpG site in the sequence is determined by the ratio of T to C signals, reflecting the proportion of unmethylated to methylated cytosine at each CpG site in the original sequence.
[0265] V. Classifiers, Machine Learning Models, and Systems
[0266] In various instances, methylation sequencing features are used as input datasets to trained algorithms (e.g., machine learning models or classifiers) to look for correlations between sequence composition and patient groupings. Examples of such patient groupings include the presence of disease or symptom, stage, subtype, responders vs. non-responders, and progressive vs. non-progressive patients. In various instances, feature matrices are generated to compare samples obtained from individuals with known conditions or characteristics. In some implementations, samples are obtained from healthy individuals or individuals without any known indications, and samples are obtained from patients known to have cancer.
[0267] As used in this article, in the context of machine learning and pattern recognition, the term "feature" generally refers to a single, measurable characteristic or feature of an observed phenomenon. The concept of "feature" relates to the concept of explanatory variables used in statistical techniques, such as, but not limited to, linear and logistic regression. Features are typically numerical, but structural features, such as strings and graphs, are used in grammatical pattern recognition.
[0268] As used herein, the term "input feature" (or "feature") generally refers to variables used by a trained algorithm (e.g., a model or classifier) to predict the output classification (label) of a sample, such as conditions, sequence content (e.g., mutations), suggested data collection actions, or suggested treatments. The value of a variable can be determined for a sample and used to determine the classification.
[0269] In each instance, input features of the genetic data include: alignment variables, which relate to the alignment of sequence data (e.g., sequence reads) with the genome, and non-alignment variables, such as those relating to the sequence content of the reads, measurements of proteins or autoantibodies, or the average methylation level of genomic regions. Input features can be genetic characteristics, such as V-plotting metrics, FREE-C unconvolution, chromatin accessibility, and cfDNA measurements at transcription start sites. Metrics that can be used in methylation analysis include, but are not limited to: the percentage of methylation of each base of CpG, CHG, and CHH; conversion efficiency (100 - average methylation percentage of CHH); low-methylated segments; methylation levels (overall average methylation of CpG, CHH, and CHG; fragment length; fragment midpoint; and methylation levels in one or more genomic regions such as chrM, LINE1, or ALU); the number of methylated CpGs per fragment; the fraction of CpG methylation per fragment relative to total CpG; the fraction of CpG methylation per region relative to total CpG; the fraction of CpG methylation per panel relative to total CpG; dinucleotide coverage (normalized dinucleotide coverage); coverage uniformity (unique CpG sites under 1x and 10x average genomic coverage (for S4 runs)); overall average CpG coverage (depth); and average coverage at CpG islands, CGI racks, and CGI shores. These metrics can be used as feature inputs for machine learning methods and models.
[0270] For multiple measurements, the system identifies a feature set to be input into a trained algorithm (e.g., a machine learning model or classifier). The system analyzes each molecular category and forms a feature vector from the measurements. The system then inputs the feature vector into the machine learning model and outputs a classification indicating whether the biological sample possesses the specified characteristic.
[0271] In some implementations, the machine learning model outputs a classifier capable of distinguishing between two or more groups or categories of individuals or features within a group of individuals or features of a group. In some implementations, the classifier is a trained machine learning classifier.
[0272] In some implementations, informative loci or characteristics of biomarkers in tumor tissue are analyzed to form a map. Receiver operating characteristic (ROC) curves can be generated by plotting the performance of a specific characteristic (e.g., any biomarker described herein and / or any additional biomedical information item) in distinguishing between two populations (e.g., individuals who respond to a treatment and those who do not). In some implementations, characteristic data across the entire population (e.g., cases and controls) are sorted in ascending order based on individual characteristic values.
[0273] In each instance, the specified characteristics are selected from health and cancer, disease subtype, disease stage, progressive and non-progressive, and responder and non-responder.
[0274] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0275] A. Data Analysis
[0276] In some instances, this disclosure provides a system, method, or kit in which data analysis can be implemented in a software application, computing hardware, or both. In various instances, the analytical application or system includes at least one data receiving module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In some embodiments, the data receiving module may include a computer system that connects laboratory hardware or instruments to a computer system that processes laboratory data. In some embodiments, the data preprocessing module may include a hardware system or computer software that performs operations on the data for analysis. Examples of operations that can be applied to the data in the preprocessing module include affine transformation, denoising, data cleaning, reformatting, or subsampling. The data analysis module may be specifically designed to analyze genomic data from one or more genomic materials; for example, it may acquire assembled genomic sequences and perform probabilistic and statistical analyses to identify anomalous patterns associated with disease, pathology, state, risk, condition, or phenotype. The data interpretation module may use analytical methods, such as those derived from statistics, mathematics, or biology, to support understanding the relationship between identified anomalous patterns and health status, functional state, prognosis, or risk. The data visualization module can use mathematical modeling, computer graphics, or rendering methods to create visual representations of data that can facilitate the understanding or interpretation of results.
[0277] In various instances, machine learning methods are applied to distinguish samples within a sample population. In some implementations, machine learning methods are applied to differentiate between healthy samples and samples with advanced diseases (such as adenomas).
[0278] In some implementations, one or more machine learning operations used to train the prediction engine include one or more of the following: generalized linear models, generalized additive models, nonparametric regression operations, random forest classifiers, spatial regression operations, Bayesian regression models, time series analysis, Bayesian networks, Gaussian networks, decision tree learning operations, artificial neural networks, recurrent neural networks, convolutional neural networks, reinforcement learning operations, linear or nonlinear regression operations, support vector machines, clustering operations, and genetic algorithm operations.
[0279] In each instance, the computer processing methods are selected from logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient-advancing trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed random neighborhood embedding (t-SNE), multilayer perceptron (MLP), network clustering, neurofuzzy logic, and artificial neural networks.
[0280] In some instances, the methods disclosed herein may include computational analysis of nucleic acid sequencing data from samples from one or more individuals.
[0281] B. Classifier Generation
[0282] On one hand, the disclosed system and method provide a classifier that is generated based on feature information derived from methylation sequence analysis of cfDNA biological samples. The classifier forms part of a prediction engine used to distinguish groups within a population based on sequence features identified in biological samples (such as cfDNA).
[0283] In some implementations, a classifier is created by the following steps: normalizing the sequence information by formatting similar portions of the sequence information into a uniform format and uniform size; storing the normalized sequence information in a columnar database; training the prediction engine by applying one or more machine learning operations to the stored normalized sequence information to map a combination of one or more features for a specific group; applying the prediction engine to accessed field information to identify individuals associated with the group; and classifying the individuals into the group.
[0284] In some implementations, the hierarchy is created by the following steps: normalizing the sequence information by formatting similar portions of the sequence information into a uniform format and uniform size; storing the normalized sequence information in a columnar database; training the prediction engine by applying one or more machine learning operations to the stored normalized sequence information to map a combination of one or more features for a specific group; applying the prediction engine to accessed field information to identify individuals related to the group; and grouping the individuals into the groups.
[0285] Specificity, as used in this article, generally refers to "the probability of a negative test result among individuals without the disease." It can be calculated by dividing the number of disease-free individuals with negative test results by the total number of disease-free individuals.
[0286] In each instance, the model, classifier, or prediction test has the following specificity: at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
[0287] As used in this article, sensitivity generally refers to "the probability of a positive test result among infected individuals." It can be calculated by dividing the number of infected individuals with positive test results by the total number of infected individuals.
[0288] In each instance, the model, classifier, or prediction test has the following sensitivities: at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
[0289] As used in this article, the positive predictive value generally refers to the "probability that a positive test result is correct". It can be calculated by dividing the number of true positive test results by the total number of positive test results.
[0290] In each instance, the model, classifier, or prediction test has the following positive predictive value: at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
[0291] The negative predictive value used in this article generally refers to the "probability that a negative test result is correct". It can be calculated by dividing the number of true negative test results by the total number of negative test results.
[0292] In each instance, the model, classifier, or prediction test has the following negative predictive value: at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
[0293] C. Digital processing device
[0294] In some instances, the subject matter described herein may include a digital processing device or its uses. In some instances, a digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform device functions. In some instances, a digital processing device may include an operating system configured to execute executable instructions.
[0295] In some instances, the digital processing device is optionally connected to a computer network. In some instances, the digital processing device is optionally connected to the Internet. In some instances, the digital processing device is optionally connected to a cloud computing facility. In some instances, the digital processing device is optionally connected to an intranet. In some instances, the digital processing device is optionally connected to a data storage device.
[0296] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, notebook computers, subnotebook computers, netbook computers, netboard computers, set-top box computers, handheld computers, internet-connected devices, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, those with brochures, notepads, and convertible configurations.
[0297] In some instances, a digital processing device may include an operating system configured to execute executable instructions. For example, an operating system may include software, including programs and data, for managing the device's hardware and providing services for the execution of applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, and OpenBSD. Linux Mac OS X Windows and Non-limiting examples of suitable personal computer operating systems include Mac OS And UNIX-like operating systems, such as In some instances, the operating system may be provided by cloud computing, and cloud computing resources may be provided by one or more service providers.
[0298] In some instances, the apparatus may include storage and / or memory devices. Storage and / or memory devices may be one or more physical devices used for temporarily or permanently storing data or programs. In some instances, the apparatus may be volatile memory and requires power to maintain the stored information. In some instances, the apparatus is non-volatile memory and retains the stored information when the digital processing apparatus is not powered. In some instances, non-volatile memory may include flash memory. In some instances, non-volatile memory may include dynamic random access memory (DRAM). In some instances, non-volatile memory may include ferroelectric random access memory (FRAM). In some instances, non-volatile memory may include phase-change random access memory (PRAM).
[0299] In some instances, the device may be a storage device, including, for example, a CD-ROM, DVD, flash memory device, disk drive, tape drive, optical disc drive, and cloud-based storage. In some instances, the storage and / or memory device may be a combination of devices such as those disclosed herein. In some instances, the digital processing device may include a display that sends visual information to a user. In some instances, the display may be a cathode ray tube (CRT). In some instances, the display may be a liquid crystal display (LCD). In some instances, the display may be a thin-film transistor liquid crystal display (TFT-LCD). In some instances, the display may be an organic light-emitting diode (OLED) display. In some instances, the OLED display may be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some instances, the display may be a plasma display. In some instances, the display may be a video projector. In some instances, the display may be a combination of devices such as those disclosed herein.
[0300] In some instances, the digital processing apparatus may include an input device for receiving information from a user. In some instances, the input device may be a keyboard. In some instances, the input device may be a pointing device, including, for example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some instances, the input device may be a touchscreen or multi-touchscreen. In some instances, the input device may be a microphone for capturing voice or other sound input. In some instances, the input device may be a camera for capturing motion or visual input. In some instances, the input device may be a combination of devices such as those disclosed herein.
[0301] D. Non-transitory computer-readable storage medium
[0302] In some instances, the subject matter disclosed herein may include one or more non-transitory computer-readable storage media encoded with a program containing instructions executable by an operating system of an optional network digital processing device. In some instances, the computer-readable storage medium may be a tangible component of a digital processing device. In some instances, the computer-readable storage medium may optionally be removable from the digital processing device. In some instances, the computer-readable storage medium may include, for example, CD-ROMs, DVDs, flash memory devices, solid-state storage, disk drives, tape drives, optical disc drives, cloud computing systems and services, etc. In some instances, programs and instructions may be permanently, substantially permanently, semi-permanently, or non-transitory encoded on the medium.
[0303] E. Computer System
[0304] This disclosure provides computer systems that are programmed to implement the methods described herein. Figure 1 A computer system 101 is shown, which is programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, and reference sequences. The computer system 101 can process various aspects of the patient data, biological data, biological sequences, or reference sequences disclosed herein. The computer system 101 can be a user's electronic device or a computer system located remotely at the electronic device. The electronic device can be a mobile electronic device.
[0305] Computer system 101 includes a central processing unit (CPU, also referred to herein as a “processor” and “computer processor”) 105, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 101 also includes memory or storage location 110 (e.g., random access memory, read-only memory, flash memory), electronic storage unit 115 (e.g., hard disk), communication interface 120 for communicating with one or more other systems (e.g., network adapter), and peripheral devices 125, such as cache, other memory, data storage, and / or electronic display adapters. Memory 110, storage unit 115, interface 120, and peripheral devices 125 communicate with CPU 105 via a communication bus (solid line) (such as a motherboard). Storage unit 115 may be a data storage unit (or data repository) for storing data. Computer system 101 may be operatively coupled to computer network (“network”) 130 by means of communication interface 120. Network 130 may be the Internet, the Internet of Things, and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some instances, network 130 is a telecommunications and / or data network. Network 130 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some instances, network 130 can be a peer-to-peer network using computer system 101, which allows devices coupled to computer system 101 to act as clients or servers.
[0306] CPU 105 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location (such as memory 110). The instructions may be directed to CPU 105, which may subsequently be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by CPU 105 may include fetching, decoding, executing, and writing back.
[0307] CPU 105 may be part of a circuit (such as an integrated circuit). One or more other components of system 101 may be included in the circuit. In some instances, the circuit is an application-specific integrated circuit (ASIC).
[0308] Storage unit 115 may store files, such as drivers, libraries, and saved programs. Storage unit 115 may store user data, such as user preferences and user programs. In some instances, computer system 101 may include one or more additional data storage units external to computer system 101, such as those located on a remote server communicating with computer system 101 via an intranet or the Internet.
[0309] Computer system 101 can communicate with one or more remote computer systems via network 130. For example, computer system 101 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate / tablet PCs (e.g., PCs with slate / tablets), etc. iPad Galaxy Tab), phone, smartphone (e.g., iPhone, Android-compatible devices (or personal digital assistant). Users can access computer system 101 via network 130.
[0310] The methods described herein can be implemented using machine-executable code (e.g., a computer processor) stored in an electronic storage location of computer system 101 (e.g., stored in memory 110 or electronic storage unit 115). The machine-executable or machine-readable code can be provided in software form. During use, the code can be executed by processor 105. In some instances, the code can be retrieved from storage unit 115 and stored in memory 110 for access by processor 105. In some instances, electronic storage unit 115 can be excluded, and the machine-executable instructions can be stored in memory 110.
[0311] Code can be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or it can be interpreted or compiled at runtime. Code can be provided in a programming language, which can be selected to enable the code to be executed in a pre-compiled, interpreted, or compiled manner.
[0312] Aspects of the systems and methods provided herein, such as computer system 101, can be embodied in programming. Various aspects of the technology can be considered "products" or "artifacts," typically in the form of machine (or processor) executable code and / or associated data, carried or contained in a type of machine-readable medium. Machine-executable code can be stored in electronic storage units such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. "Storage" type media can include any or all tangible memory or related modules of computers, processors, etc., such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable the loading of software from one computer or processor into another computer or processor, e.g., from a management server or host computer into a computer platform for an application server. Therefore, another type of medium that can carry software elements includes optical, electrical, and electromagnetic waves, such as those used on physical interfaces between local devices via wired and optical terrestrial networks and various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.
[0313] Therefore, machine-readable media (such as computer-executable code) can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, any storage device such as any one or more computers, such as those that can be used to implement the database shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wires and optical fibers, including wires that form the bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, floppy hard disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card tapes, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or cassette tapes, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other media from which a computer may read programming code and / or data. Many of these forms of computer-readable media may involve transmitting one or more sequences of one or more instructions to a processor for execution.
[0314] Computer system 101 may include or communicate with an electronic display 135, the electronic display 135 including a user interface (UI) 140 for providing, for example, analysis of nucleic acid sequences, concentrated nucleic acid samples, methylation profiles, expression profiles, and methylation or expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0315] The methods and systems disclosed herein can be implemented using one or more algorithms. These algorithms can be implemented in software when executed by the central processing unit 105. For example, the algorithms can store, process, identify, or interpret patient data, biological data, biological sequences, and reference sequences.
[0316] While certain examples of methods and systems have been shown and described herein, those skilled in the art will recognize that these are provided by way of example only and are not intended to be limiting. Many variations, modifications, and substitutions will appear to those skilled in the art without departing from the scope of the description herein. Furthermore, it should be understood that all aspects of the methods and systems described are not limited to the specific descriptions, configurations, or relative proportions listed herein, which depend on a variety of conditions and variables, and the descriptions are intended to include such alternatives, modifications, variations, or equivalents.
[0317] In some instances, the subject matter disclosed herein may include at least one computer program or its purpose. A computer program may be a sequence of instructions written to perform a specified task and executed in a CPU, GPU, or TPU of a digital processing device. Computer-readable instructions may be implemented as program modules that perform a specific task or implement a specific abstract data type, such as functions, objects, application programming interfaces (APIs), data structures, etc. Given the disclosure provided herein, computer programs can be written in various versions of various languages.
[0318] In various environments, the functionality of computer-readable instructions can be combined or assigned as needed. In some instances, a computer program may include a single sequence of instructions. In some instances, a computer program may include multiple sequences of instructions. In some instances, a computer program may be provided from one location. In some instances, a computer program may be provided from multiple locations. In some instances, a computer program may include one or more software modules. In some instances, a computer program may include, in whole or in part, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ons, or add-ons, or combinations thereof.
[0319] In some instances, computer processing can be methods from statistics, mathematics, biology, or any combination thereof. In some instances, computer processing methods include dimensionality reduction methods, such as logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis functions, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient-driven trees, logistic regression, matrix factorization, network clustering, and neural networks.
[0320] In some instances, the computer processing method is a supervised machine learning method, including, for example, regression, support vector machines, tree-based methods, and networks.
[0321] In some instances, the computer processing method is an unsupervised machine learning approach, including, for example, clustering, networks, principal component analysis, and matrix factorization.
[0322] F. Database
[0323] In some instances, the subject matter disclosed herein may include one or more databases, or the use of such databases to store patient data, biological data, biological sequences, or reference sequences. Reference sequences may be derived from the database. Given the disclosures provided herein, many databases are suitable for storing and retrieving sequence information. In some instances, suitable databases may include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relational model databases, association databases, and XML databases. In some instances, the database may be internet-based. In some instances, the database may be network-based. In some instances, the database may be cloud-based. In some instances, the database may be based on one or more local computer storage devices.
[0324] On the one hand, this disclosure provides a non-transitory computer-readable medium including instructions that instruct a processor to perform the methods described herein.
[0325] On the one hand, this disclosure provides a computing device including a computer-readable medium.
[0326] On the other hand, this disclosure provides a system for classifying biological samples, comprising: a) a receiver for receiving a plurality of training samples, each of the plurality of training samples having molecules of a plurality of classes, wherein each of the plurality of training samples contains one or more known labels; b) a feature module for identifying operable feature sets corresponding to measurements for inputting into a machine learning model for each of the plurality of training samples, wherein the feature sets correspond to molecular characteristics in the plurality of training samples, wherein for each of the plurality of training samples, the system is operable to perform a plurality of different measurements on molecules of a plurality of classes in the training samples to obtain a set of measurements, wherein each set of measurements is derived from a single measurement performed on one class of molecules in the training samples, wherein a plurality of sets of measurements are obtained for the plurality of training samples; c) an analysis module for analyzing the sets of measurements. The analysis is performed to obtain training vectors for the training samples, wherein the training vectors include feature values of N feature sets corresponding to the measurements, each feature value corresponding to a feature and including one or more measurements, wherein the training vectors are formed using at least two features from at least two of the N feature sets corresponding to a first subset of multiple different measurements; d) a labeling module for informing the system about the training vectors using the parameters of the machine learning model in order to obtain output labels for multiple training samples; e) a comparator module for comparing the output labels with known labels of the training samples; f) a training module for iteratively searching for the optimal values of the parameters as part of training the machine learning model, which is based on the comparison of the output labels with the known labels of the training samples; and g) an output module for providing the parameters of the machine learning model and the feature set of the machine learning model.
[0327] VI. Methods for classifying objects within a group
[0328] The disclosed method aims to determine genetic and / or epigenetic parameters of genomic DNA associated with colonic proliferative disorders by analyzing cfDNA in subjects. The method is intended to improve the diagnosis, treatment, and surveillance of colonic proliferative disorders, and more specifically, to improve the identification and differentiation between stages or subtypes of the disorder and their genetic susceptibility.
[0329] In some implementations, the method includes analyzing the methylation status of CpG islands, CpG shores, or CpG skeletons.
[0330] In some implementations, the method includes analyzing the methylation status, hemimethylation status, hypermethylation status, or hypomethylation status of cell-free nucleic acids in biological samples.
[0331] On one hand, this disclosure provides a method for detecting colonic proliferative disorder (CPD), which can be applied to cell-free samples, for example, to detect CPD DNA in cell-free circulating samples. The method utilizes the detection of methylation signals in a single sequencing read as a fundamental "positive" CPD signal.
[0332] In some implementations, colonic proliferative disorders are selected from: adenomas (adenomatous polyps), sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas. In some implementations, colonic proliferative disorders include colorectal cancer.
[0333] On one hand, this disclosure provides a method for detecting colonic proliferative disease, comprising: extracting DNA from a cell-free sample obtained from a subject; converting at least a portion of the DNA for methylation sequencing; amplifying methylated regions in the cancer generated by the converted DNA; generating sequencing reads from the amplified regions; and detecting a signal of colonic proliferative disease comprising at least one, at least two, at least three, or more than three methylated regions within a cancer panel to obtain input features, which are fed into a machine learning model to obtain a classifier capable of distinguishing between two groups of subjects (e.g., healthy vs. cancer, disease stage, advanced adenoma vs. cancer).
[0334] The trained machine learning methods, models, and discriminative classifiers described in this paper can be applied to a variety of medical applications, including cancer detection, diagnosis, and treatment responsiveness. Because the models can be trained using individual metadata and analyte-derived features, applications can be customized to stratify individuals within a population and guide treatment decisions accordingly.
[0335] diagnosis
[0336] The methods and systems provided herein can perform predictive analytics using artificial intelligence-based approaches to analyze data obtained from a subject (patient) and generate diagnostic output for the subject with cancer (e.g., colorectal cancer). For example, the application can apply predictive algorithms to the acquired data to generate a diagnosis for the subject with cancer. The predictive algorithms can include artificial intelligence-based predictors, such as machine learning-based predictors, configured to process the acquired data to generate a diagnosis for the subject with cancer.
[0337] Machine learning predictors can be trained using datasets, for example, using a dataset generated by methylation assays of individual biological samples from one or more cohorts of cancer patients using the signature panel described herein as input, and the known diagnostic results (e.g., stage and / or tumor score) of the objects as outputs of the machine learning predictor.
[0338] Training datasets (e.g., datasets generated from methylation assays of individual biological samples using the signature panel described herein) can be generated from one or more sets of objects, for example, that share common characteristics (features) and outcomes (tags). Training datasets may include sets of features and tags corresponding to diagnostically relevant features. Features may include characteristics such as certain ranges or categories measured by cfDNA assays, such as the counts of cfDNA fragments in sets of bins (genome windows) that overlap or fall within a reference genome from biological samples obtained from healthy and diseased samples. For example, a set of features collected from a given object at a given time point can collectively serve as a diagnostic signature, indicating that the object has an identified cancer at that time point. Features may also include tags indicating the object's diagnostic outcome, such as one or more cancers.
[0339] The tags can include results, such as the subject's known diagnosis (e.g., stage and / or tumor score). Results can include cancer-related characteristics of the subject. For example, a characteristic could indicate that the subject has one or more types of cancer.
[0340] The training set (e.g., the training dataset) can be selected by random sampling of a dataset corresponding to one or more subject sets (e.g., a retrospective and / or prospective cohort of patients with or without one or more cancers). Alternatively, the training set (e.g., the training dataset) can be selected by proportional sampling of a dataset corresponding to one or more subject sets (e.g., a retrospective and / or prospective cohort of patients with or without one or more cancers). The training set can be balanced among datasets corresponding to one or more subject sets (e.g., patients from different clinical sites or trials). The machine learning predictor can be trained until certain predetermined accuracy or performance conditions are met, such as having a minimum expected value corresponding to a diagnostic accuracy metric. For example, the diagnostic accuracy metric might correspond to a prediction of the diagnosis, stage, or tumor score of one or more cancers in a subject.
[0341] Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and the area under the receiver operating characteristic (ROC) curve (AUC) corresponding to the diagnostic accuracy in detecting or predicting cancer (e.g., colorectal cancer).
[0342] On one hand, this disclosure provides a method for using a classifier capable of distinguishing individual groups, comprising: a) analyzing molecules of multiple categories in a biological sample, wherein the analysis provides multiple sets of measurements representing the multiple categories of molecules; b) identifying a set of features corresponding to a characteristic of each of the multiple categories of molecules input into a machine learning or statistical model; c) preparing a feature vector of feature values from each of the multiple sets of measurements, each feature value corresponding to a feature of the feature set and including one or more measurements, wherein the feature vector includes at least one feature value obtained using each of the multiple sets of measurements; d) loading the following into the memory of a computer system: a trained machine learning model including a classifier, a trained machine learning model trained using training vectors obtained from training biological samples, a first subset of training biological samples identified as having a specified characteristic, and a second subset of training biological samples identified as not having the specified characteristic; and e) applying the trained machine learning model to the feature vectors to obtain an output classification of whether the biological sample has the specified characteristic, thereby distinguishing individual groups having the specified characteristic.
[0343] On one hand, this disclosure provides a method using a hierarchical structure capable of distinguishing individual groups, comprising: a) analyzing molecules of multiple categories in a biological sample, wherein the analysis provides multiple sets of measurements representing the multiple categories of molecules; b) identifying a set of features corresponding to a characteristic of each of the multiple categories of molecules input into a machine learning or statistical model; c) preparing a feature vector of feature values from each of the multiple sets of measurements, each feature value corresponding to a feature of the feature set and including one or more measurements, wherein the feature vector includes at least one feature value obtained using each of the multiple sets of measurements; d) loading the following into the memory of a computer system: a trained machine learning model including a classifier, a trained machine learning model trained using training vectors obtained from training biological samples, a first subset of training biological samples identified as having a specified characteristic, and a second subset of training biological samples identified as not having the specified characteristic; and e) applying the trained machine learning model to the feature vectors to obtain an output classification of whether the biological sample has the specified characteristic, thereby distinguishing individual groups having the specified characteristic.
[0344] On one hand, this disclosure provides a method for using a hierarchical structure capable of distinguishing individual groups, comprising: a) detecting methylation signals in a single sequencing read of a preselected genomic region in one or more first patient samples, b) the methylation signals affecting the hierarchical structure of the data output, thereby affecting a machine learning model, and c) using the affected hierarchical structure to detect methylation signals in a second patient sample.
[0345] In some implementations, the pre-selected genomic regions are selected from two or more methylated genomic regions in Tables 1-11, three or more methylated genomic regions in Tables 1-11, four or more methylated genomic regions in Tables 1-11, five or more methylated genomic regions in Tables 1-11, six or more methylated genomic regions in Tables 1-11, seven or more methylated genomic regions in Tables 1-11, eight or more methylated genomic regions in Tables 1-11, nine or more methylated genomic regions in Tables 1-11, ten or more methylated genomic regions in Tables 1-11, eleven or more methylated genomic regions in Tables 1-11, twelve or more methylated genomic regions in Tables 1-11, or thirteen or more methylated genomic regions in Tables 1-11.
[0346] In another aspect, this disclosure provides a method for identifying cancer in a subject, comprising: a) providing a biological sample from the subject containing cell-free nucleic acid (cfNA) molecules; b) methylating and sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; c) aligning the plurality of cfNA sequencing reads to a reference genome; d) generating quantitative measures of the plurality of cfNA sequencing reads on each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprise at least about 10 distinct regions, each of the at least about 10 distinct regions comprising at least a portion of a gene selected from methylated regions in the signature panel described herein; and e) applying a trained algorithm to the first cfNA feature set to generate the probability that the subject has the cancer.
[0347] In some instances, the at least about 10 distinct regions comprise at least about 20 distinct regions, each of the at least about 20 distinct regions comprising at least a portion of the methylated regions identified in Tables 1-11. In some instances, the at least about 10 distinct regions comprise at least about 30 distinct regions, each of the at least about 30 distinct regions comprising at least a portion of the methylated regions identified in Tables 1-11.
[0348] As another example, such predetermined conditions could include values that predict the specificity of colonic proliferative disorders: for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0349] As another example, such predetermined conditions could be positive predictive value (PPV) for predicting proliferative colorectal disease, including values such as: at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0350] As another example, such predetermined conditions could be negative predictive value (NPV) for predicting colonic proliferative disorders, including values such as: at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0351] As another example, such predetermined conditions could be the area under the receiver operating characteristic (ROC) curve (AUC) of a predictor of colonic proliferative disease, including the following values: at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0352] Treatment responsiveness
[0353] The predictive classifiers, systems, and methods described herein can be used to classify populations of individuals for a variety of clinical applications (e.g., based on methylation assays of individual biological samples using the signature panel described herein). Examples of such clinical applications include: detecting early-stage cancer, diagnosing cancer, classifying cancer into specific disease stages, and determining responsiveness or resistance to therapeutic agents used to treat cancer.
[0354] The methods and systems described herein can be applied to the characteristics of colonic proliferative disorders, such as grading and stage. Therefore, combinations of analytes and assays can be used in this system and method to predict the responsiveness of cancer treatments to different cancer types in different tissues and to classify individuals based on treatment responsiveness. In some embodiments, the classifier described herein is capable of stratifying a group of individuals into treatment responders and non-responders.
[0355] This disclosure also provides a method for identifying drug targets (e.g., genes associated with or important to a specific category) for a target symptom or disease, comprising: assessing the gene expression level of at least one gene in a sample obtained from an individual; and using a neighborhood analysis procedure to identify genes associated with the sample classification, thereby identifying one or more classification-related drug targets.
[0356] This disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, comprising: obtaining a sample from an individual suffering from said disease class; subjecting the sample to drug treatment; assessing the gene expression level of at least one gene in the drug-exposed sample; and classifying the drug-exposed sample into a disease class based on a computer model established using a weighted voting scheme, according to a function of the sample’s relative gene expression level with respect to the model.
[0357] This disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, wherein an individual has been treated with the drug, the method comprising obtaining a sample from the individual treated with the drug; assessing the gene expression level of at least one gene in the sample; and classifying the sample into a disease class using a model established using a weighted voting scheme, including assessing the gene expression level of the sample compared with the gene expression level of the model.
[0358] This disclosure also provides a method for determining whether an individual belongs to a phenotypic category (e.g., intelligence, response to treatment, lifespan, likelihood of viral infection, or obesity), comprising: obtaining a sample from the individual; assessing the gene expression level of at least one gene in the sample; and classifying the sample into a disease class using a model established using a weighted voting scheme, including assessing the gene expression level of the sample compared to the gene expression level of the model.
[0359] On the one hand, the systems and methods described herein related to population classification based on treatment response refer to, but are not limited to, cancers treated with chemotherapeutic agents (including, but not limited to, DNA-damaging agents, DNA repair-targeted therapies, DNA damage signaling inhibitors, DNA damage-induced cell cycle arrest inhibitors, and inhibitors of processes that indirectly cause DNA damage). Each of these chemotherapeutic agents can be considered a “DNA damage therapeutic agent,” as used in this document.
[0360] Based on patient analyte data, patients can be grouped into high-risk and low-risk groups, such as those with a high or low clinical risk of recurrence, and the results can be used to determine the treatment course. For example, patients identified as high-risk may receive adjuvant chemotherapy after surgery. For patients considered low-risk, adjuvant chemotherapy may be discontinued after surgery. Therefore, this disclosure provides, in some aspects, a method for preparing a colorectal cancer tumor gene expression profile that indicates the risk of recurrence.
[0361] In each instance, the classifier described in this paper is able to stratify individual groups between those who respond to treatment and those who do not.
[0362] On the other hand, the methods disclosed in this article can be applied to clinical applications involving cancer detection or monitoring.
[0363] In some implementations, the methods disclosed herein can be used to determine and / or predict responses to treatment.
[0364] In some implementations, the methods disclosed herein can be used to monitor and / or predict tumor burden.
[0365] In some implementations, the methods disclosed herein can be used to detect and / or predict postoperative residual tumor.
[0366] In some implementations, the methods disclosed herein can be used to detect and / or predict minimal residual disease after treatment.
[0367] In some implementations, the methods disclosed herein can be used to detect and / or predict recurrence.
[0368] On the one hand, the method disclosed in this paper can be used for secondary screening.
[0369] On the one hand, the method disclosed in this paper can be used as a screening tool.
[0370] On the one hand, the methods disclosed in this article can be used to monitor cancer development.
[0371] On the one hand, the methods disclosed in this article can be used to monitor and / or predict cancer risk.
[0372] VII. Identification or monitoring of colorectal cancer
[0373] After processing a dataset using a trained algorithm, colorectal cancer can be identified or monitored in objects. The identification can be based, at least in part, on quantitative measures of sequence reads from a dataset of colorectal cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA from colorectal cancer-associated genomic loci).
[0374] Colorectal cancer can be identified in subjects with the following accuracy: at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The accuracy of colorectal cancer identification by the trained algorithm can be calculated as the percentage of independently tested samples (e.g., subjects known to have colorectal cancer or subjects with negative clinical test results for colorectal cancer) correctly identified or classified as having or not having colorectal cancer.
[0375] Colorectal cancer can be identified in subjects with the following positive predictive value (PPV): at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The percentage of colorectal cancer identified using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as having colorectal cancer compared to those actually having colorectal cancer.
[0376] Colorectal cancer can be identified in subjects with the following negative predictive value (NPV): at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. The NPV for identifying colorectal cancer using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as not having colorectal cancer compared to those actually having colorectal cancer.
[0377] Colorectal cancer can be identified in subjects with the following clinical sensitivities: at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%. At least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or greater. The clinical sensitivity for identifying colorectal cancer using a trained algorithm can be calculated as the percentage of independently tested samples associated with the presence of colorectal cancer (e.g., subjects known to have colorectal cancer) that are correctly identified or classified as having colorectal cancer.
[0378] Colorectal cancer can be identified in subjects with the following clinical specificity: at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%. At least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or greater. The clinical specificity for identifying colorectal cancer using a trained algorithm can be calculated as the percentage of independent test samples (e.g., those with negative clinical test results for colorectal cancer) that are correctly identified or classified as not having colorectal cancer.
[0379] In some implementations, the trained algorithm can determine that a subject has a risk of developing colorectal cancer of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater.
[0380] The trained algorithm can determine whether an individual has a risk of colorectal cancer with an accuracy of at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, and at least approximately 90%. 2%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater.
[0381] After a subject is diagnosed with colorectal cancer, a therapeutic intervention can be provided (e.g., prescribing or administering an appropriate treatment course for colorectal cancer). A therapeutic intervention may include prescribing an effective dose of medication, further testing or evaluation for colorectal cancer, further monitoring for colorectal cancer, or a combination thereof. If the subject is currently receiving treatment for colorectal cancer with a course of therapy, a therapeutic intervention may include a subsequent different course of therapy (e.g., to increase efficacy if the current course of therapy is ineffective). Therapeutic interventions may be described, for example, by reference to “WHO list of priority medical devices for cancer management, WHO Medical device technical series”, World Health Organization, ISBN: 978-92-4-156546-2, Geneva, 2017, the contents of which are incorporated herein by reference. Therapeutic interventions can be described, for example, by Wolpin et al., “Systemic Treatment of Colorectal Cancer,” Gastroenterology, Vol. 134, No. 5, 2008, pp. 1296-1310.e1, the contents of which are incorporated herein by reference.
[0382] Therapeutic interventions may include recommending a second clinical examination to confirm the diagnosis of colorectal cancer. This second clinical examination may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0383] Quantitative measures of sequence reads from a colorectal cancer-related genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-related genomic loci) can be evaluated over a period of time to monitor patients (e.g., those with colorectal cancer or undergoing colorectal cancer treatment). In this context, the quantitative measures of the patient dataset can change during treatment. For example, the quantitative measures of a patient dataset whose risk of colorectal cancer has decreased due to effective treatment can shift to the profile or distribution of healthy subjects (e.g., those without colorectal cancer). Conversely, the quantitative measures of a patient dataset whose risk of colorectal cancer has increased due to ineffective treatment can shift to the profile or distribution of subjects with a higher risk of colorectal cancer or a higher grade or stage of colorectal cancer.
[0384] The colorectal cancer in a subject can be monitored by monitoring the treatment process. Monitoring may include assessing the subject's colorectal cancer at two or more time points. The assessment may be based at least on quantitative measures of sequence reads from datasets on colorectal cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including quantitative measures of colorectal cancer-associated genomic loci identified at each of the two or more time points.
[0385] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic loci (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including differences in quantitative measures of colorectal cancer-associated genomic loci determined between two or more time points, can indicate one or more clinical indications, such as: (i) a colorectal cancer diagnosis in the subject; (ii) a colorectal cancer prognosis in the subject; (iii) an increased risk of colorectal cancer in the subject; (iv) a decreased risk of colorectal cancer in the subject; (v) the efficacy of a treatment course for colorectal cancer in the subject; and (vi) the ineffectiveness of a treatment course for colorectal cancer in the subject.
[0386] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic loci (CLC) datasets (e.g., quantitative measures of RNA transcripts or DNA at CLC genomic loci), including differences in quantitative measures of CLC genomic loci determined between two or more time points, can indicate a diagnosis of colorectal cancer in a subject. For example, if colorectal cancer was not detected in a subject at an earlier time point but was detected at a later time point, then the difference indicates a diagnosis of colorectal cancer. Clinical actions or decisions can be made based on this indication of a colorectal cancer diagnosis in the subject, for example, prescribing or administering a new therapeutic intervention. Clinical actions or decisions may include recommending a second clinical examination to confirm the diagnosis of colorectal cancer. This secondary clinical examination may include imaging examinations, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0387] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic locus panels (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including differences in quantitative measures of colorectal cancer-associated genomic locus panels determined between two or more time points, can indicate the prognosis of the target colorectal cancer.
[0388] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic loci (CLC) datasets (e.g., quantitative measures of RNA transcripts or DNA at CLC loci), including differences in quantitative measures of CLC loci determined between two or more time points, can indicate an increased risk of colorectal cancer in a subject. For example, if a subject is diagnosed with colorectal cancer at both an earlier and later time point, and if the difference is positive (e.g., the quantitative measures of sequence reads from CLC loci datasets (e.g., quantitative measures of RNA transcripts or DNA at CLC loci) increase from an earlier to a later time point), then the difference can indicate an increased risk of colorectal cancer in the subject. Clinical actions or decisions can be made based on this indication of increased risk of colorectal cancer, such as prescribing or administering a new therapeutic intervention or switching a therapeutic intervention (e.g., ending current treatment and prescribing or administering a new treatment). Clinical action or decisions may include recommending a second clinical examination to confirm an increased risk of colorectal cancer. This second clinical examination may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0389] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic loci (CLC) datasets (e.g., quantitative measures of RNA transcripts or DNA at CLC loci), including differences in quantitative measures of CLC loci determined between two or more time points, can indicate a reduced risk of colorectal cancer in the subject. For example, if a subject is diagnosed with colorectal cancer at both an earlier and later time point, and if the difference is negative (e.g., quantitative measures of sequence reads from CLC loci datasets (e.g., quantitative measures of RNA transcripts or DNA at CLC loci), including quantitative measures of CLC loci, decrease from an earlier to a later time point), then the difference can indicate a reduced risk of colorectal cancer in the subject. Clinical actions or decisions can be made based on this indication of reduced risk of colorectal cancer for the subject (e.g., continuing or ending current therapeutic interventions). Clinical action or decisions may include recommending a second clinical examination to confirm a reduced risk of colorectal cancer. This second clinical examination may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0390] In some implementations, differences in quantitative measures of sequence reads from colorectal cancer-associated genomic loci (CMA) datasets (e.g., quantitative measures of RNA transcripts or DNA at CMA loci), including differences in quantitative measures of CMA loci determined between two or more time points, can indicate the efficacy of a treatment course for colorectal cancer in a subject. For example, if colorectal cancer was detected in a subject at an earlier time point but not at a later time point, the difference can indicate the efficacy of the treatment course for colorectal cancer in the subject. Clinical actions or decisions can be made based on this indication of the efficacy of the treatment course for colorectal cancer in a subject, for example, continuing or discontinuing the current therapeutic intervention for the subject. Clinical actions or decisions may include recommending a second clinical examination for the subject to confirm the efficacy of the treatment course for colorectal cancer. This secondary clinical examination may include imaging examinations, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0391] In some implementations, differences in quantitative measures of sequence reads from a dataset on a colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including differences in quantitative measures of colorectal cancer-associated genomic locus panels determined between two or more time points, can indicate that the treatment process for colorectal cancer in a subject is ineffective. For example, if colorectal cancer is detected in a subject at both an earlier and a later time point, and if the difference is positive or zero (e.g., quantitative measures of sequence reads from a dataset on a colorectal cancer-associated genomic locus panel (e.g., quantitative measures of RNA transcripts or DNA at colorectal cancer-associated genomic loci), including quantitative measures of colorectal cancer-associated genomic locus panels, increase or remain at a constant level from the earlier to the later time point), then the difference can indicate that the treatment process for colorectal cancer in a subject is ineffective. Clinical actions or decisions may be made based on the indication that the treatment course for colorectal cancer is ineffective, for example, ending the current therapeutic intervention and / or switching (e.g., prescribing or administering) a new, different therapeutic intervention. Clinical actions or decisions may include recommending a second clinical examination to confirm the ineffectiveness of the treatment course for colorectal cancer. This second clinical examination may include imaging studies, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, fecal immunochemical testing (FIT), fecal occult blood testing (FOBT), or any combination thereof.
[0392] VIII. Reagent Kit
[0393] This disclosure provides a kit for identifying or monitoring cancer in a subject. The kit may include probes for identifying sequence quantification measures (e.g., indicating presence, absence, or relative quantity) at each of a plurality of cancer-related genomic loci in a cell-free biological sample of the subject. The sequence quantification measures (e.g., indicating presence, absence, or relative quantity) at each of the plurality of cancer-related genomic loci in the cell-free biological sample may indicate one or more types of cancer. The probes may be selective for sequences at the plurality of cancer-related genomic loci in the cell-free biological sample. The kit may include instructions for treating a cell-free biological sample with the probes to generate a dataset indicating sequence quantification measures (e.g., indicating presence, absence, or relative quantity) at each of the plurality of cancer-related genomic loci in the cell-free biological sample of the subject.
[0394] The probes in the kit are selective for sequences at multiple cancer-associated genomic loci in cell-free biological samples. The probes in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can be sequence complementary to nucleic acid sequences from one or more of multiple cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more different cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions can include one or more members selected from the regions listed in Tables 1-11.
[0395] Instructions for use in the kit may include instructions for determining the cell-free biological sample using probes selective for sequences at multiple cancer-associated genomic loci. These probes may be nucleic acid molecules (e.g., RNA or DNA) that are sequence-complementary to nucleic acid sequences (e.g., RNA or DNA) from one or more of the multiple cancer-associated genomic loci. These nucleic acid molecules may be primers or enriched sequences. Instructions for determining the cell-free biological sample may include instructions for performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample, thereby generating a dataset indicating a quantitative measure of sequence at each of the multiple cancer-associated genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative quantity). This quantitative measure of sequence at each of the multiple cancer-associated genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative quantity) may indicate one or more cancers.
[0396] The kit's instructions may include instructions for measuring and interpreting assay readouts that can be quantified at one or more of multiple cancer-associated genomic loci to generate a dataset indicating quantitative measures of sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample (e.g., indicating presence, absence, or relative quantity). For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple cancer-associated genomic loci can generate a dataset indicating quantitative measures of sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample (e.g., indicating presence, absence, or relative quantity). Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof.
[0397] This invention provides, but is not limited to, the following embodiments:
[0398] 1. A methylated signature panel specific to colonic proliferative disorders, comprising:
[0399] One or more methylated genomic regions selected from Table 11, wherein the methylation level of the one or more regions is higher in biological samples from individuals with colonic proliferative disease or a subtype of colonic proliferative disease, and lower in normal tissues and normal blood cells from individuals without colonic proliferative disease.
[0400] 2. The methylation signature panel as described in Embodiment 1, wherein the biological sample is nucleic acid, DNA, RNA or cell-free nucleic acid (cfDNA or cfRNA).
[0401] 3. A methylation signature panel as described in Embodiment 1, wherein the signature panel contains increased methylation in two or more genomic regions selected from Table 11.
[0402] 4. The methylated signature panel as described in Embodiment 1, wherein the colonic proliferative disease is selected from: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0403] 5. The methylated signature panel as described in Embodiment 1, wherein the colonic cell proliferative disease is selected from stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.
[0404] 6. A methylation signature panel as described in Embodiment 1, wherein the signature panel comprises two or more methylated genomic regions in Tables 1-11, three or more methylated genomic regions in Tables 1-11, four or more methylated genomic regions in Tables 1-11, five or more methylated genomic regions in Tables 1-11, six or more methylated genomic regions in Tables 1-11, seven or more methylated genomic regions in Tables 1-11, eight or more methylated genomic regions in Tables 1-11, nine or more methylated genomic regions in Tables 1-11, ten or more methylated genomic regions in Tables 1-11, eleven or more methylated genomic regions in Tables 1-11, twelve or more methylated genomic regions in Tables 1-11, or thirteen or more methylated genomic regions in Tables 1-11.
[0405] 7. A methylated signature panel as described in Embodiment 1, wherein the signature panel contains methylated genomic regions in colorectal cancer, including methylated regions selected from one or more of the following genomic regions: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1.
[0406] 8. The methylation signature panel as described in Embodiment 1, wherein the methylated regions in colorectal cancer include methylated regions selected from the following genomic regions: IKZF1, KCNQ5, and ELMO1.
[0407] 9. The methylation signature panel as described in Embodiment 1, wherein the methylated regions in colorectal cancer include methylated regions selected from one or more of the following genomic regions: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1.
[0408] 10. A methylated signature panel as described in Embodiment 1, wherein the signature panel includes methylated genomic regions selected from Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11.
[0409] 11. A methylated signature panel specific to colonic proliferative disorders, comprising:
[0410] Two or more methylated genomic regions selected from Tables 1-11, wherein the two or more regions are more methylated in biological samples from individuals with colonic proliferative disease or a subtype of colonic proliferative disease, and less methylated in normal tissues and normal blood cells from individuals without colonic proliferative disease.
[0411] 12. A methylation signature panel as described in Embodiment 11, wherein the biological sample is nucleic acid, DNA, RNA or cell-free nucleic acid.
[0412] 13. A methylation signature panel as described in Embodiment 11, wherein the signature panel contains methylation enhancements in six or more genomic regions selected from Tables 1-11.
[0413] 14. The methylated signature panel as described in Embodiment 11, wherein the colonic proliferative disease is selected from: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0414] 15. The methylated signature panel as described in Embodiment 11, wherein the colonic cell proliferative disease is selected from stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.
[0415] 16. A methylation signature panel as described in Embodiment 11, wherein the signature panel comprises three or more methylated genomic regions in Tables 1-11, four or more methylated genomic regions in Tables 1-11, five or more methylated genomic regions in Tables 1-11, six or more methylated genomic regions in Tables 1-11, seven or more methylated genomic regions in Tables 1-11, eight or more methylated genomic regions in Tables 1-11, nine or more methylated genomic regions in Tables 1-11, ten or more methylated genomic regions in Tables 1-11, eleven or more methylated genomic regions in Tables 1-11, twelve or more methylated genomic regions in Tables 1-11, or thirteen or more methylated genomic regions in Tables 1-11.
[0416] 17. A methylated signature panel as described in Embodiment 11, wherein the signature panel contains methylated genomic regions in colorectal cancer, including methylated regions selected from one or more of the following genomic regions: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, and FLI1.
[0417] 18. The methylation signature panel as described in Embodiment 11, wherein the methylated regions in colorectal cancer include methylated regions selected from the following genomic regions: IKZF1, KCNQ5, and ELMO1.
[0418] 19. The methylation signature panel as described in Embodiment 11, wherein the methylated regions in colorectal cancer include methylated regions selected from one or more of the following genomic regions: IKZF1, KCNQ5, ELMO1, CHST2, PRKCB, FLI1, CLIP4, ELOVL5, FAM72B, and ST3GAL1.
[0419] 20. A methylated signature panel as described in Embodiment 11, wherein the signature panel includes methylated genomic regions selected from Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11.
[0420] 21. A machine learning classifier capable of distinguishing between a group of healthy individuals and a group of individuals with colonic cell proliferative disorder, comprising:
[0421] a) A set of measurements representing differentially methylated genomic regions as described in Embodiment 1, wherein the measurements are obtained from methylation sequencing data from healthy subjects and subjects with colonic cell proliferation disease;
[0422] b) wherein the measured values are used to generate a feature set corresponding to the characteristics of the differentially methylated genomic regions, and wherein the features are input into a machine learning or statistical model;
[0423] c) Wherein the model provides feature vectors that can be used as a classifier to distinguish between a group of healthy individuals and a group of individuals with colonic cell proliferation disorder.
[0424] 22. A classifier as described in embodiment 21, wherein the set of measurements describes characteristics of methylated regions selected from: the percentage of base-by-base methylation of CpG, CHG, and CHH; the count or ratio of fragments with different counts or ratios of methylated CpG observed in the region; conversion efficiency (100 - average methylation percentage of CHH); low-methylated segments; methylation level (overall average methylation of CpG, CHH, and CHG; fragment length; fragment midpoint; number of methylated CpGs per fragment; fraction of CpG methylation per fragment of total CpG; fraction of CpG methylation per region of total CpG; fraction of CpG methylation per panel of total CpG; dinucleotide coverage (normalized dinucleotide coverage); coverage uniformity (unique CpG sites under 1x and 10x average genome coverage (for S4 run)); overall average CpG coverage (depth); and average coverage at CpG islands, CGI racks, and CGI shores.
[0425] 23. A system for detecting colonic cell proliferative disorders, comprising a machine learning model classifier, wherein:
[0426] a) A computer-readable medium including a classifier operable to classify objects according to a methylation signature panel as having or not having colonic proliferative disease; and
[0427] (b) One or more processors for executing instructions stored on the computer-readable medium.
[0428] 24. The system of embodiment 23, comprising a classifier of embodiment 21 loaded into the memory of a computer system, a machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having colonic proliferative disease, and a second subset of the training biological samples identified as not having colonic proliferative disease.
[0429] 25. A method for determining the methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample from an individual, comprising:
[0430] a) Provide conditions that enable the conversion of unmethylated cytosine to uracil in the nucleic acid molecules of the cfDNA sample to produce multiple transformed nucleic acids;
[0431] b) Contact the plurality of transformed nucleic acids with a nucleic acid probe, the nucleic acid probe being complementary to a pre-identified methylation signature panel selected from at least two differentially methylated regions in Tables 1-11, to enrich sequences corresponding to the signature panel;
[0432] c) Determine the nucleic acid sequences of the plurality of transformed nucleic acid molecules; and
[0433] d) Align the nucleic acid sequences of the plurality of transformed nucleic acid molecules with a reference nucleic acid sequence to determine the methylation profile of the individual.
[0434] 26. The method of embodiment 25 further includes amplifying the plurality of transformed nucleic acids.
[0435] 27. The method of embodiment 26, wherein the amplification includes polymerase chain reaction (PCR).
[0436] 28. The method of embodiment 25 further includes determining the nucleic acid sequence of the transformed nucleic acid molecule at a depth greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x.
[0437] 29. The method of embodiment 25, wherein the reference nucleic acid sequence is at least a portion of the human reference genome.
[0438] 30. The method as described in embodiment 29, wherein the human reference genome is hg18.
[0439] 31. The method of embodiment 25, wherein the pre-identified methylation signature panel comprises three or more methylated genomic regions in Tables 1-11, four or more methylated genomic regions in Tables 1-11, five or more methylated genomic regions in Tables 1-11, six or more methylated genomic regions in Tables 1-11, seven or more methylated genomic regions in Tables 1-11, eight or more methylated genomic regions in Tables 1-11, nine or more methylated genomic regions in Tables 1-11, ten or more methylated genomic regions in Tables 1-11, eleven or more methylated genomic regions in Tables 1-11, twelve or more methylated genomic regions in Tables 1-11, or thirteen or more methylated genomic regions in Tables 1-11.
[0440] 32. The method of embodiment 31, wherein the pre-identified methylation signature panel comprises one or more methylated genomic regions in Table 11, two or more methylated genomic regions in Table 11, or three methylated genomic regions in Table 11.
[0441] 33. The method of embodiment 25, wherein the methylation profile indicates the presence or absence of colonic cell proliferative disease in the individual.
[0442] 34. The method of embodiment 33, wherein the colonic proliferative disease is selected from: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0443] 35. The method of embodiment 33, wherein the colonic cell proliferative disease is selected from stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, or stage 4 colorectal cancer.
[0444] 36. A method for detecting the presence or absence of colonic cell proliferative disorder in a subject, comprising:
[0445] a) Provides conditions for converting unmethylated cytosine into uracil in nucleic acid molecules of biological samples obtained from or derived from the object to produce multiple converted nucleic acids;
[0446] b) Contact the plurality of transformed nucleic acids with a nucleic acid probe, the nucleic acid probe being complementary to a pre-identified methylation signature panel selected from at least two differentially methylated regions in Tables 1-11, to enrich sequences corresponding to the signature panel;
[0447] c) Determine the nucleic acid sequence of the transformed nucleic acid molecule;
[0448] d) Align the nucleic acid sequences of the plurality of transformed nucleic acid molecules with a reference nucleic acid sequence to determine the methylation profile of the individual; and
[0449] e) Applying a trained machine learning classifier to the methylation spectrum, wherein the trained machine learning classifier is trained to distinguish between healthy individuals and individuals with colonic proliferative disorder, to provide output values associated with the presence of colonic proliferative disorder, thereby detecting the presence or absence of colonic proliferative disorder in the subject.
[0450] 37. The method of embodiment 36, wherein the biological sample obtained from the object is selected from: cell-free DNA, cell-free RNA, body fluids, feces, colon excretions, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0451] 38. The method of embodiment 36 further includes amplifying the plurality of transformed nucleic acids.
[0452] 39. The method of embodiment 38, wherein the amplification comprises polymerase chain reaction (PCR).
[0453] 40. The method of embodiment 36 further includes determining the nucleic acid sequence of the transformed nucleic acid molecule at a depth greater than 1000x, greater than 2000x, greater than 3000x, greater than 4000x, or greater than 5000x.
[0454] 41. The method of embodiment 36, wherein the reference nucleic acid sequence is at least a portion of the human reference genome.
[0455] 42. The method as described in embodiment 41, wherein the human reference genome is hg18.
[0456] 43. The method of embodiment 36, wherein the pre-identified methylation signature panel comprises three or more methylated genomic regions in Tables 1-11, four or more methylated genomic regions in Tables 1-11, five or more methylated genomic regions in Tables 1-11, six or more methylated genomic regions in Tables 1-11, seven or more methylated genomic regions in Tables 1-11, eight or more methylated genomic regions in Tables 1-11, nine or more methylated genomic regions in Tables 1-11, ten or more methylated genomic regions in Tables 1-11, eleven or more methylated genomic regions in Tables 1-11, twelve or more methylated genomic regions in Tables 1-11, or thirteen or more methylated genomic regions in Tables 1-11.
[0457] 44. The method of embodiment 43, wherein the pre-identified methylation signature panel comprises one or more methylated genomic regions in Table 11, two or more methylated genomic regions in Table 11, or three methylated genomic regions in Table 11.
[0458] 45. The method of embodiment 36 further includes administering treatment for the colonic cell proliferative disorder to the individual based on the detection of the presence of the colonic cell proliferative disorder in the individual.
[0459] 46. The method of embodiment 36, wherein the colonic cell proliferative disease is selected from: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0460] 47. The method of embodiment 36, wherein the colonic proliferative disease includes colorectal cancer.
[0461] 48. The method of embodiment 36, wherein the colonic cell proliferative disease is selected from stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer and stage 4 colorectal cancer.
[0462] 49. The method of embodiment 36, wherein the trained machine learning classifier is selected from: deep learning classifiers, neural network classifiers, linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, linear kernel support vector machine classifiers, first-order or second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic net algorithm classifiers, sequence minimum optimization algorithm classifiers, Naive Bayes algorithm classifiers, and principal component analysis classifiers.
[0463] Example
[0464] Example 1: Selection of methylated regions for colorectal cancer detection
[0465] For colorectal cancer, using the systems and methods of this disclosure, 20 highly methylated genomic regions were identified in the tumor, but many of these regions were not methylated in normal tissues. These regions were used as highly specific markers of the presence of tumors with little or no background signal.
[0466] In Table 12, 'Start-End Position' specifies the coordinates of the target region in the construction of the human genome reference sequence hg18. Gene IDs and chromosome fields refer to the gene and chromosome numbers associated with the numbered regions. Examination of these sequences relative to neighboring genes indicates they are found upstream, in 5' promoters, 5' enhancers, introns, exons, distal promoters, coding regions, or intergenic regions, respectively.
[0467] use Cell-free DNA isolation kit (Applied) Cell-free DNA (doped with a unique synthetic double-stranded DNA (dsDNA) fragment for sample tracking) was extracted from 250 microliters (μL) of plasma according to the manufacturer's instructions.
[0468] UltraII DNA Library Preparation Kit (New England) Prepare paired-end sequencing libraries, including polymerase chain reaction (PCR) amplification and unique molecular identifiers (UMIs), and use... The NovaSeq 6000 sequencing system sequences at 2x5 l base pairs on multiple S2 or S4 stream cells up to a minimum of 400 million reads (median = 636 million reads).
[0469] probes targeting colorectal cancer
[0470] PCR primer pairs were developed for different regions of the genome that showed extensive methylation in multiple colorectal cancer samples from the TOGA database, but little or no methylation in multiple normal tissues and blood cells (peripheral blood mononuclear cells and others).
[0471] These primers were then used to amplify transformed DNA from plasma samples from individuals at risk of colorectal cancer. Sequencing aptamers were ligated to the DNA, and next-generation sequencing was performed. The sequencing reads were then separated by region and analyzed using tools such as the BiQ Analyzer HT program.
[0472] The obtained sequencing reads were demultiplexed, aptamer-trimmed, and aligned with the human reference genome (GRCh38 with decoys, alt contigs, and HLA contigs) using a Burrows Wheeler alignment machine (BWA-MEM 0.7.15). PCR replicates, if present, were removed using fragment endpoints and / or UMI.
[0473] A cfDNA “profile” is created for each sample by counting the number of fragments aligned with each putative protein-coding region in the genome. This type of data visualization reveals epigenetic changes in cfDNA protected by variable nucleosomes, resulting in observed changes in fragment coverage and methylation compared to controls.
[0474] The set of functional regions of the human genome, including putative protein-coding gene regions (genomic coordinates including introns and exons), is annotated in the sequencing data. The annotation of protein-coding gene regions (“gene” regions) is derived from the Integrated Human Expressed Sequences (CHESS) project (v1.0).
[0475] The results are as follows.
[0476] Table 12 provides a collection of highly methylated genomic regions in cell-free nucleic acid samples identified from individuals with colorectal cancer. For each region, an exemplary number of methylated CpG sites in that region is provided as a threshold for distinguishing healthy individuals from CRC individuals.
[0477] Table 12
[0478]
[0479]
[0480] In this discussion, references to genes such as ITGA4, TMEM163, and SMBT2 may not indicate the gene of interest itself, but rather the relevant methylation region described in the signature panel.
[0481] A total of 50 regions were found to be hypermethylated and associated with CRC. Not all regions needed to be included in the classification model to distinguish between healthy and CRC individuals. Therefore, some regions appeared to broadly indicate the various types of cancer being assessed. Other regions were methylated in these subgroups, while the remainder were cancer-specific. In the context of this assay and the types of cancer examined, certain regions could be described as “particularly methylated in colorectal cancer” and given higher weights in the signature when training sample sequences in the prediction model. These CRC-associated, high-weighted methylated regions were used in a specific model that…
[0482] Training was conducted to differentiate between healthy individuals and CRC individuals.
[0483] Example 2: Constructing and training a classification model to distinguish individual populations of colorectal cancer
[0484] Using the systems and methods disclosed herein, machine learning classification models are built and trained using artificial intelligence-based approaches to analyze cfDNA data obtained from objects (generating diagnostic output for objects with colorectal cancer).
[0485] The anticipated human plasma samples were obtained from 49 patients diagnosed with CRC. Additionally, a collection of 92 control samples was obtained from patients who currently have no cancer diagnosis (but may have other comorbidities or undiagnosed cancers). All samples were de-identified.
[0486] For each sample, obtain the age, sex, and cancer stage (if available) of each patient. Plasma samples collected from each patient will be stored at -80°C and thawed before use. Table 13 provides a description of the study cohort, showing the number of healthy and cancer samples used for the CRC experiments (by stage, sex, and age).
[0487] Table 13
[0488]
[0489] Samples were processed and sequenced according to the methods described herein, particularly those in Example 1. The methylated regions in Table 12 are specifically used to determine the methylated CpG status between healthy individuals and those with colorectal cancer. For each region listed in column 1 of Table 12, the threshold number of CpG sites shown in column 2 was used to define the methylated fragments available for analysis. Other fragments were classified as methylated if they had multiple CpG sites greater than the threshold; otherwise, they were classified as unmethylated. To calculate the raw score for each sample, these counts for each sample were aggregated across regions, the raw score being the number of methylated fragments in each sample overlapping the regions listed in Table 12. The raw score for each sample was normalized to account for coverage differences between samples. The raw score for each sample was multiplied by a sample-specific scaling factor, given by dividing the total number of samples by a pre-specified target coverage level. These normalized and scaled methylation ratios output as the score for each sample. A threshold score was selected based on the desired specific targets from the training set. Based on whether the scores of these samples exceed this threshold, the samples are classified as positive or negative. An ROC curve is generated by considering the grades of samples with this score or the threshold.
[0490] The machine learning classification model was trained as described above, and parameters were selected on a separately proposed sample set. The machine learning classification model was applied to the samples described in Table 13. Healthy samples with the highest proportion of hypermethylated fragment counts were selected as the cutoff value for classifying new samples as positive or negative. The area under the ROC curve (AUC) was calculated based on the training set described above, using the rank obtained from the normalized hypermethylated fragment counts. Sensitivity and specificity were calculated using the selected cutoff value. Confidence intervals for sensitivity and specificity were calculated using Clopper-Pearson confidence intervals, and confidence intervals for AUC were calculated using the method described in Fay, M. and Malinovsky, Y., Statistics in Medicine 37(27):3991-4006(2018) (the contents of which are incorporated herein by reference).
[0491] The mean area under the curve (AUC) of this method was 0.9488 (0.87–0.98), and the mean sensitivity for IU samples was 70% (0.49–0.87) with 92% specificity (0.86–0.96). Figure 2 ).
[0492] Example 3: Detection and individual classification of cell-free samples
[0493] Using the systems and methods of this disclosure, predictive analysis is performed using an artificial intelligence-based approach to analyze cfDNA data obtained from an object, thereby generating a diagnostic output for an object suffering from colorectal cancer.
[0494] This article provides a method for predicting an increased risk of developing or progressing to cancer in asymptomatic patients, wherein a model trained on a signature panel of the procedure provided in Example 1 is applied to a panel of measured biomarkers, and clinical factors of age and sex are used to identify those patients with an increased risk of developing or progressing to colorectal cancer. In an embodiment, this method and the classifier model of the present invention use biomarkers measured within the normal clinical range as input variables, wherein when the output of a first classifier model is higher than a calculated threshold based on the number of methylated CpG sites in the region, the colorectal cancer classifier model uses the age input variable and measurements from the patient's biomarker panel to classify the patient into an increased risk category.
[0495] According to Example 1, the purpose of selecting gene markers and CpG sites is to select those with strong differential methylation (β-difference, i.e., the difference and p-value between methylation-specific probes and methylation-non-specific probes), predictive ability (AUC), and influence on gene expression (from the p-value of gene expression).
[0496] This selection resulted in the signature panel presented herein, containing methylated regions that can distinguish healthy samples from CRC samples. The first subset of regions contains 20 regions with at least 4 to 18 CpG sites of increased methylation, which map to 18 genes (many genes are represented by many CpG sites).
[0497] The cfDNA CpG counting profile of input cfDNA can serve as an unbiased representation of available methylation signals in the blood, allowing the capture of signals directly from the tumor as well as those from non-tumor sources such as the circulating immune system or the tumor microenvironment.
[0498] Unsupervised clustering based on these genes revealed clear methylation patterns associated with healthy or CRC phenotypes.
[0499] To evaluate the accuracy of methylated regions for early detection of CRC, receiver operating characteristic (ROC) curves and area under the ROC curve (AUC) of the regions in the signature panel were calculated. Figures 3A-3F The ROC results are shown, demonstrating the ability of these differentially methylated regions (DMRs) to detect CRC and differentiate early-stage cancers, including those with stage 1 ( Figure 3A Phase 2 Figure 3B Phase 3 Figure 3C ), 4th phase ( Figure 3D ), missing phase ( Figure 3E ) and all samples ( Figure 3F Patients with CRC were included. A total of 80 gene regions associated with increased methylation were identified. Methylation regions with a gradually increasing average methylation level compared to controls could be used to differentiate between early and late CRC. For example, the methylation regions associated with Table 12 had a high CRC detection capability [CRC vs. Control AUC = 0.924 (95% CI: 0.752 to 0.954)].
[0500] As summarized in Table 14, the results demonstrate excellent performance in early cancer detection from blood (e.g., in a set of 13 stage I and II samples).
[0501] Table 14
[0502]
[0503]
[0504] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited to the specific embodiments provided herein. Although the invention has been described with reference to the foregoing detailed description, the description and illustrative examples of embodiments herein are not intended to be construed as limiting. Various variations, modifications, and alternatives will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein can be used to implement the invention. Therefore, it is contemplated that the invention should also cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the invention and are thereby intended to cover the methods and structures within the scope of these claims and their equivalents.
Claims
1. A methylation signature panel specific for a colon cell proliferative disorder comprising: one or more methylation genomic regions selected from Table 11, wherein the one or more regions are more highly methylated in a biological sample from an individual having a colon cell proliferative disorder or a subtype of a colon cell proliferative disorder and less methylated in normal tissue and normal blood cells from an individual not having a colon cell proliferative disorder.
2. The methylation signature panel of claim 1, wherein the biological sample is nucleic acid, DNA, RNA, or cell free nucleic acid (cfDNA or cfRNA).
3. The methylation signature panel of claim 1, wherein the signature panel comprises increased methylation in two or more genomic regions selected from Table 11.
4. The methylation signature panel of claim 1, wherein the colon cell proliferative disorder is selected from the group consisting of: adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal carcinoma, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
5. The methylation signature panel of claim 1, wherein the colon cell proliferative disorder is selected from the group consisting of stage 1 colorectal cancer, stage 2 colorectal cancer, stage 3 colorectal cancer, and stage 4 colorectal cancer.
6. A methylation signature panel specific for a colon cell proliferative disorder comprising: two or more methylation genomic regions selected from Tables 1-11, wherein the two or more regions are more highly methylated in a biological sample from an individual having a colon cell proliferative disorder or a subtype of a colon cell proliferative disorder and less methylated in normal tissue and normal blood cells from an individual not having a colon cell proliferative disorder.
7. A machine learning classifier capable of distinguishing between a population of healthy individuals and a population of individuals having a colon cell proliferative disorder comprising: a) a set of measurements representing the differentially methylated genomic regions of claim 1, wherein the measurements are obtained from methylation sequencing data from healthy subjects and subjects having a colon cell proliferative disorder; b) wherein the measurements are used to generate a set of features corresponding to the properties of the differentially methylated genomic regions, and wherein the features are input to a machine learning or statistical model; c) wherein the model provides a feature vector that is used as a classifier capable of distinguishing between a population of healthy individuals and a population of individuals having a colon cell proliferative disorder.
8. A system for detecting a colon cell proliferative disorder comprising a machine learning model classifier comprising: a) a computer readable medium comprising a classifier operable to classify a subject as having a colon cell proliferative disorder or not having a colon cell proliferative disorder according to a methylation signature panel; and b) one or more processors for executing instructions stored on the computer readable medium.
9. A method for determining a methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample from an individual, comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules of the cfDNA sample to produce a plurality of converted nucleic acids; b) contacting the plurality of converted nucleic acids with nucleic acid probes complementary to a pre-identified methylation signature panel of at least two differentially methylated regions selected from Tables 1-11 to enrich for sequences corresponding to the signature panel; c) determining nucleic acid sequences of the plurality of converted nucleic acid molecules; and d) aligning the nucleic acid sequences of the plurality of converted nucleic acid molecules to reference nucleic acid sequences, thereby determining the methylation profile of the individual.
10. A method for detecting the presence or absence of a colon cell proliferative disorder in a subject, comprising: a) providing conditions capable of converting unmethylated cytosines to uracils in nucleic acid molecules of a biological sample obtained or derived from the subject to produce a plurality of converted nucleic acids; b) contacting the plurality of converted nucleic acids with nucleic acid probes complementary to a pre-identified methylation signature panel of at least two differentially methylated regions selected from Tables 1-11 to enrich for sequences corresponding to the signature panel; c) determining nucleic acid sequences of the converted nucleic acid molecules; d) aligning the nucleic acid sequences of the plurality of converted nucleic acid molecules to reference nucleic acid sequences, thereby determining the methylation profile of the individual; and e) applying a trained machine learning classifier to the methylation profile, wherein the trained machine learning classifier is trained to be capable of distinguishing between healthy individuals and individuals having a colon cell proliferative disorder to provide an output value correlating to the presence of a colon cell proliferative disorder, thereby detecting the presence or absence of the colon cell proliferative disorder in the subject.
Citation Information
Patent Citations
Method of detection of methylated nucleic acid using agents which modify unmethylated cytosine and distinguishing modified methylated and non-methylated nucleic acids
US5786146A