Method and system for processing cell-free nucleic acids
The method processes cfDNA samples to generate methylation profiles, addressing the challenges of low sensitivity and specificity in ctDNA detection by enriching methylated nucleic acids, achieving high accuracy in diagnosing cancers with AUROC values up to 99%.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ADELA INC
- Filing Date
- 2024-04-12
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for detecting cancer using circulating tumor DNA (ctDNA) face challenges in sensitivity and specificity, particularly in cases where ctDNA is present in low amounts, making it difficult to accurately diagnose low-shedding or early-stage cancers.
A method involving the processing of cell-free deoxyribonucleotide (cfDNA) samples to generate methylation profiles through sequencing and computer processing, utilizing methylation profiles to determine the presence of cancer with high accuracy, including the use of methylation-binding molecules and capture reagents to enrich methylated nucleic acids, followed by sequencing without bisulfite conversion.
The method achieves high sensitivity and specificity in detecting various cancers, with AUROC values ranging from 89% to 99%, enabling early-stage and low-shedding cancer detection with improved accuracy.
Smart Images

Figure 2026515793000001_ABST
Abstract
Description
Technical Field
[0001]
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 496,347, filed Apr. 14, 2023; U.S. Provisional Patent Application No. 63 / 501,359, filed May 10, 2023; U.S. Provisional Patent Application No. 63 / 511,441, filed Jun. 30, 2023; U.S. Provisional Patent Application No. 63 / 517,327, filed Aug. 2, 2023; U.S. Provisional Patent Application No. 63 / 588,120, filed Oct. 5, 2023; U.S. Provisional Patent Application No. 63 / 591,732, filed Oct. 19, 2023; U.S. Provisional Patent Application No. 63 / 594,365, filed Oct. 30, 2023; U.S. Provisional Patent Application No. 63 / 602,156, filed Nov. 22, 2023; U.S. Provisional Patent Application No. 63 / 549,294, filed Feb. 2, 2024; and U.S. Provisional Patent Application No. 63 / 571,139, filed Mar. 28, 2024, each of which is hereby incorporated by reference in its entirety.
Background Art
[0002]
[0002] Circulating tumor DNA (ctDNA) is increasingly showing potential as a non-invasive tumor-specific biomarker for routine clinical use. CtDNA is mainly derived from tumor cells undergoing cell death and is released into circulation in various body fluids, including blood. In most cancer patients, the majority of cell-free DNA in blood is derived from healthy (e.g., non-cancerous) tissue. Furthermore, the observed proportion of ctDNA can range from <0.1% to 90% of total cell-free DNA at diagnosis, depending on several factors including the primary site of the tumor and disease burden. CtDNA provides non-invasive access to the molecular status and disease burden of the tumor.
[0003] Incorporation by Reference
[0003] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Overview of the project] [Means for solving the problem]
[0004]
[0004] In some embodiments, the method provided herein includes the steps of (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleotide (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer with at least about 91% receiver operating characteristic area (AUROC), wherein the cancer is a low-shedding cancer. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from a low-shedding tumor. In some embodiments, the cancer is bladder cancer, breast cancer, endometrial cancer, prostate cancer, or kidney cancer. In some embodiments, the cancer is endometrial cancer or prostate cancer.
[0005]
[0005] In some embodiments, the method provided herein includes the steps of (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleotide (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 94%, wherein the cancer is an early-stage cancer. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from an early-stage tumor. In some embodiments, the early-stage tumor is a stage I tumor. In some embodiments, the early-stage tumor is a stage II tumor. In some embodiments, the cancer is bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, prostate cancer, or kidney cancer. In some embodiments, the cancer is esophageal cancer, hepatobiliary cancer, or ovarian cancer.
[0006]
[0006] In some embodiments, the method provided herein includes the steps of (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleotide (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads in the absence of bisulfite conversion; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer, wherein the cancer is endometrial cancer, esophageal cancer, hepatobiliary cancer, ovarian cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer. In some embodiments, the cfDNA sample comprises a circulating tumor nucleic acid molecule derived from endometrial cancer. In some embodiments, the methylation profiles are computer processed to determine that the subject has endometrial cancer with at least about 90% receiver operating characteristic area (AUROC). In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from esophageal cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has esophageal cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from hepatobiliary tract cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has hepatobiliary tract cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from ovarian cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has ovarian cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 97%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from prostate cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has prostate cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 89%.In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from bladder cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has bladder cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 95%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from breast cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has breast cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 92%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from colorectal cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has colorectal cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 98%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from head and neck cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has head and neck cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 96%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from lung cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has lung cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 96%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from pancreatic cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has pancreatic cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from kidney cancer. In some embodiments, the methylation profile is computer-processed to determine that the subject has kidney cancer with an area under the receiver operating characteristic curve (AUROC) of at least about 91%.
[0007]
[0007] In some embodiments, the method further includes, prior to (b), adding a set of nucleic acid molecules not derived from the subject to the plurality of nucleic acid molecules. In some embodiments, the methylation profile is genome-wide. In some embodiments, the methylation profile includes the whole methylome. In some embodiments, the method (e.g., step (d)) includes a supervised machine learning method, the supervised machine learning method being regression, support vector machines, tree-based methods, neural networks, or nearest neighbor methods. In some embodiments, the method (e.g., step (d)) includes an unsupervised machine learning method, the unsupervised machine learning method being clustering, neural networks, principal component analysis, or matrix factorization. In some embodiments, the subject has previously been treated for cancer and the cancer has substantially disappeared, and (d) includes the step of determining that the subject has a recurrence of the cancer. In some embodiments, the method further includes, prior to (b), adding a certain amount of filler DNA to the plurality of nucleic acid molecules or derivatives thereof. In some embodiments, the amount of filler DNA includes double-stranded DNA. In some embodiments, the amount of filler DNA is about 20 nanograms (ng) to about 100 ng. In some embodiments, at least a portion of the filler DNA is methylated. In some embodiments, 10% to 40% of the filler DNA is methylated, and the remainder is unmethylated filler DNA. In some embodiments, the method further comprises, before (b), contacting the cfDNA sample with a methylated nucleic acid capture reagent to generate the plurality of nucleic acid molecules, the plurality of nucleic acids comprising one or more methylated regions. In some embodiments, the methylated nucleic acid capture reagent comprises a binder and a solid substrate. In some embodiments, the methylated nucleic acid capture reagent is produced by coupling the binder to the solid substrate by incubating the binder with the solid substrate. In some embodiments, the coupling of the binder to the solid substrate is performed before the step of contacting the cfDNA sample with the methylated nucleic acid capture reagent. In some embodiments, the solid substrate is beads.In some embodiments, the solid substrate is protein A beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the binder comprises an antibody. In some embodiments, the binder is selected from the group consisting of anti-5-methylcytosine antibody or derivatives thereof, anti-5-carboxylcytosine antibody or derivatives thereof, anti-5-formylcytosine antibody or derivatives thereof, anti-5-hydroxymethylcytosine antibody or derivatives thereof, anti-3-methylcytosine antibody or derivatives thereof, and any combination thereof. In some embodiments, one or more methylated regions are enriched in the plurality of nucleic acids with at least about 99% specificity. In some embodiments, the method further comprises a step of amplifying the plurality of nucleic acid molecules to produce an amplicon before (b), and (c) comprising a step of sequencing the amplicon. In some embodiments, the amplification step is performed on the plurality of nucleic acids while the plurality of nucleic acids are bound to a solid support. In some embodiments, the amplification comprises PCR amplification. In some embodiments, the PCR amplification comprises at least 13 or at least 14 cycles. In some embodiments, the method further includes, prior to (b), a step of enriching one or more target sequences by contacting the plurality of nucleic acid molecules with one or more nucleic acid capture probes. In some embodiments, the one or more target sequences include one or more genes.
[0008]
[0008] In some embodiments, a method for processing a nucleic acid sample from a subject is provided herein, comprising the steps of: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the subject, wherein the plurality of filler nucleic acids comprises at least one methylated nucleic acid molecule; (b) (i) incubating the methylation-binding molecules with (ii) a solid substrate to form a methylated nucleic acid capture reagent; and (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture to concentrate the plurality of methylated nucleic acids in the nucleic acid sample mixture. In some embodiments, the method further comprises the step of amplifying the captured methylated nucleic acids after the capture step to generate an amplicon of the plurality of methylated nucleic acids. In some embodiments, the amplification step is carried out while the plurality of methylated nucleic acids are bound to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some embodiments, the method further comprises subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction. In some embodiments, the sequencing reaction is sequencing by a synthesis reaction. In some embodiments, the sequencing reaction does not involve bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids include cell-free nucleic acids. In some embodiments, the solid substrate is beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the solid substrate contains protein A. In some embodiments, the solid substrate contains streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule contains biotin. In some embodiments, the methylation-binding molecule is bound to methylated cytosine. In some embodiments, the method further includes, before (a), the steps of obtaining the nucleic acid sample from the subject and carrying out one or more library preparation reactions on the nucleic acid sample.In some embodiments, the method further includes the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids after the step of carrying out the one or more library preparation reactions and before (a). In some embodiments, the method further includes the step of subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids, after the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes the step of carrying out additional magnetic capture. In some embodiments, the method further includes the step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (c). In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99% specificity. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99.5% specificity. In some embodiments, the method further includes the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences include one or more genes.
[0009]
[0009] In some embodiments, a method for processing a nucleic acid sample from a subject is provided herein, comprising the steps of (a) providing a nucleic acid sample comprising a plurality of methylated nucleic acids; (b) (i) incubating methylation-binding molecules with (ii) a solid substrate to form a methylated nucleic acid capture reagent; and (c) capturing the plurality of methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, thereby concentrating the plurality of methylated nucleic acids in the nucleic acid sample mixture, wherein the plurality of methylated nucleic acids are concentrated with a specificity greater than 99%. In some embodiments, the method further comprises, after the capture step, amplifying the captured methylated nucleic acids to produce amplicons of the plurality of methylated nucleic acids. In some embodiments, the amplification step is performed while the plurality of methylated nucleic acids are bound to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some embodiments, the method further comprises subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction. In some embodiments, the sequencing reaction is sequencing by a synthesis reaction. In some embodiments, the sequencing reaction does not involve bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids include cell-free nucleic acids. In some embodiments, the solid substrate is beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the solid substrate contains protein A. In some embodiments, the solid substrate contains streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule contains biotin. In some embodiments, the methylation-binding molecule is bound to methylated cytosine. In some embodiments, the method further includes, prior to (c), a step of carrying out one or more library preparation reactions on the methylated nucleic acids.In some embodiments, the method further includes incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids after the step of carrying out the one or more library preparation reactions and before (c). In some embodiments, the method further includes subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids, after the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes carrying out additional magnetic capture. In some embodiments, the method further includes denaturing the nucleic acids in the nucleic acid sample after (a) and before (c). In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences include one or more genes.
[0010]
[0010] In some embodiments, a method for processing a nucleic acid sample from a subject is provided herein, comprising the steps of: (a) providing a nucleic acid sample comprising a plurality of methylated nucleic acids; (b) incubating (i) methylation-binding molecules with (ii) a solid substrate to form a methylated nucleic acid capture reagent; (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating methylated nucleic acids bound to a solid substrate; and (d) amplifying the methylated nucleic acids bound to the solid substrate to generate amplicons of the plurality of methylated nucleic acids. In some embodiments, the amplification step is performed while the plurality of methylated nucleic acids are bound to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some embodiments, the method further comprises subjecting the amplicons of methylated nucleic acids to a sequencing reaction. In some embodiments, the sequencing reaction is sequencing by a synthesis reaction. In some embodiments, the sequencing reaction does not involve bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids include cell-free nucleic acids. In some embodiments, the solid substrate is beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the solid substrate contains protein A. In some embodiments, the solid substrate contains streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule contains biotin. In some embodiments, the methylation-binding molecule binds to methylated cytosine. In some embodiments, the method further includes, before (c), a step of performing one or more library preparation reactions on the methylated nucleic acids. In some embodiments, the method further includes, after the step of performing one or more library preparation reactions and before (c), a step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acids.In some embodiments, the method further includes the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, followed by subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with the nucleic acid. In some embodiments, the method further includes the step of performing additional magnetic capture. In some embodiments, the method further includes the step of denaturing the nucleic acid in the nucleic acid sample after (a) and before (c). In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are concentrated at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are concentrated at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are concentrated with at least 99% specificity. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are concentrated with at least 99.5% specificity. In some embodiments, the method further includes the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to concentrate one or more target sequences. In some embodiments, the one or more target sequences include one or more genes.
[0011]
[0011] In some embodiments, a method for processing a nucleic acid sample from a subject is provided herein, comprising the steps of: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the subject, wherein the plurality of filler nucleic acids comprises at least one methylated nucleic acid molecule; (b) capturing the methylated nucleic acids by adding a capture reagent comprising a solid substrate to the nucleic acid sample mixture, thereby generating methylated nucleic acids bound to the solid substrate; and (c) amplifying the methylated nucleic acids bound to the solid substrate to generate amplicons of methylated nucleic acids. In some embodiments, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some embodiments, the method further comprises subjecting the amplicons of methylated nucleic acids to a sequencing reaction. In some embodiments, the sequencing reaction is sequencing by a synthesis reaction. In some embodiments, the sequencing reaction does not involve bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids comprises cell-free nucleic acids. In some embodiments, the solid substrate is beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the solid substrate contains protein A. In some embodiments, the solid substrate contains streptavidin. In some embodiments, the capture reagent is a methylated nucleic acid capture agent. In some embodiments, the methylated nucleic acid capture agent contains a methylated binding molecule attached to the solid substrate. In some embodiments, the methylated binding molecule is an antibody. In some embodiments, the methylated binding molecule contains biotin. In some embodiments, the methylated binding molecule is bound to methylated cytosine. In some embodiments, the method further comprises, before (a), the steps of obtaining the nucleic acid sample from the subject and carrying out one or more library preparation reactions on the nucleic acid. In some embodiments, the method further comprises, after carrying out one or more library preparation reactions and before (a), the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid.In some embodiments, the method further includes, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with the nucleic acid. In some embodiments, the method further includes performing additional magnetic capture. In some embodiments, the method further includes, after (a) and before (b), denaturing the nucleic acids in the nucleic acid sample mixture. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99% specificity. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to concentrate one or more target sequences. In some embodiments, the one or more target sequences include one or more genes.
[0012]
[0012] In some embodiments, a method is provided herein that includes (a) obtaining a first nucleic acid molecule from a cell-free sample of interest; (b) generating a second set of nucleic acid molecules from the first set of nucleic acid molecules or derivatives thereof, wherein the methylation level in the second set of nucleic acid molecules is concentrated compared to that of the first set of nucleic acid molecules; (c) concentrating one or more targets in the second set of nucleic acid molecules or derivatives thereof to obtain a third set of nucleic acid molecules; and (d) sequencing the third set of nucleic acid molecules or derivatives thereof. In some embodiments, the concentrating step includes contacting the second set of nucleic acid molecules or derivatives thereof with one or more nucleic acid capture probes. In some embodiments, the generating step includes contacting the first set of nucleic acid molecules or derivatives thereof with a methylated nucleic acid capture reagent. In some embodiments, the methylated nucleic acid capture reagent is formed by incubating methylated molecules with a solid substrate. In some embodiments, the solid substrate is beads. In some embodiments, the solid substrate is a magnetic solid substrate. In some embodiments, the solid substrate contains protein A. In some embodiments, the solid substrate comprises streptavidin. In some embodiments, the methylated molecule is an antibody. In some embodiments, the methylated molecule comprises biotin. In some embodiments, the methylated molecule is bound to methylated cytosine. In some embodiments, the method further comprises a step of amplifying the second set of molecules. In some embodiments, the amplification step is performed while a subset of the first set of nucleic acids is bound to the methylated nucleic acid capture reagent. In some embodiments, the sequencing is sequencing by a synthetic reaction. In some embodiments, the sequencing does not involve bisulfite sequencing. In some embodiments, prior to the sequencing, one or more library preparation reactions are performed on the third set of nucleic acid molecules.In some embodiments, the method includes incubating the third set of nucleic acid molecules with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out the one or more library preparation reactions and before the sequencing step. In some embodiments, the method further includes subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids, after the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes carrying out additional magnetic capture. In some embodiments, the method further includes adding a certain amount of filler DNA to the first nucleic acid molecules before (b).
[0013]
[0013] These and other features of preferred embodiments of the present invention will become more apparent in the following detailed description with reference to the accompanying drawings. [Brief explanation of the drawing]
[0014] [Figure 1]
[0014] This figure illustrates a process for collecting flow-through of unmethylated / low-methylated DNA fragments. [Figure 2]
[0015] This figure shows a schematic diagram of a computer system according to an embodiment of the disclosure. [Figure 3]
[0016] This figure shows receiver operating characteristic (ROC) curves for the entire cohort of 12 different cancers. [Figure 4-1]
[0017] Figures 4A-4L show the early detection of multiple cancers according to the type of cancer. Figure 4A shows the ROC curve (95% confidence interval) for bladder cancer, showing the area under the ROC curve (AUC) for all stages, stages I, II, III, and IV. Figure 4B shows the ROC curve (95% confidence interval) for breast cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4C shows the ROC curve (95% confidence interval) for colorectal cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4D shows the ROC curve (95% confidence interval) for endometrial cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4E shows the ROC curve (95% confidence interval) for esophageal cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4F shows the ROC curves (95% confidence interval) for head and neck cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4G shows the ROC curves (95% confidence interval) for hepatobiliary tract cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4H shows the ROC curves (95% confidence interval) for lung cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4I shows the ROC curves (95% confidence interval) for ovarian cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4J shows the ROC curves (95% confidence interval) for pancreatic cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4K shows the ROC curves (95% confidence interval) for prostate cancer, showing the AUC for all stages, stages I, II, III, and IV.Figure 4L shows the ROC curves (95% confidence intervals) for kidney cancer, illustrating the AUC for all stages, stage I, stage II, stage III, and stage IV. [Figure 4-2]Figures 4A-4L show the early detection of multiple cancers according to the type of cancer. Figure 4A shows the ROC curve (95% confidence interval) for bladder cancer, showing the area under the ROC curve (AUC) for all stages, stages I, II, III, and IV. Figure 4B shows the ROC curve (95% confidence interval) for breast cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4C shows the ROC curve (95% confidence interval) for colorectal cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4D shows the ROC curve (95% confidence interval) for endometrial cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4E shows the ROC curve (95% confidence interval) for esophageal cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4F shows the ROC curves (95% confidence interval) for head and neck cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4G shows the ROC curves (95% confidence interval) for hepatobiliary tract cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4H shows the ROC curves (95% confidence interval) for lung cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4I shows the ROC curves (95% confidence interval) for ovarian cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4J shows the ROC curves (95% confidence interval) for pancreatic cancer, showing the AUC for all stages, stages I, II, III, and IV. Figure 4K shows the ROC curves (95% confidence interval) for prostate cancer, showing the AUC for all stages, stages I, II, III, and IV.Figure 4L shows the ROC curves (95% confidence intervals) for kidney cancer, illustrating the AUC for all stages, stage I, stage II, stage III, and stage IV. [Figure 5]
[0018] Figure 5A shows ctDNA quantification and prognosis prediction in renal cancer. Figure 5B shows ctDNA quantification and prognosis prediction in renal cancer. Figure 5A shows ctDNA quantification scores for renal cancer, generated based on mean normalized counts of reduced-size fragments and adjusted for methylation specificity to identify cancer-associated methylation across 2027 regions. The ctDNA quantification score threshold was set so that 95% of samples without recurrence or progression were below the threshold. Figure 5B shows Kaplan-Meier plots of recurrence / progression-free survival probabilities for renal cancer over time. [Figure 6]
[0019] Figure 6A shows ctDNA quantification and prognosis prediction in head and neck cancer. Figure 6B shows ctDNA quantification and prognosis prediction in head and neck cancer. Figure 6A shows ctDNA quantification scores for head and neck cancer, generated based on mean normalized counts of reduced-size fragments and adjusted for methylation specificity to identify cancer-associated methylation across 2027 regions. The ctDNA quantification score threshold was set so that 95% of samples without recurrence or progression were below the threshold. Figure 6B shows Kaplan-Meier plots of recurrence / progression-free survival probabilities for head and neck cancer over time. [Figure 7]
[0020] This figure shows Kaplan-Meier plots indicating the event-free survival time over time for individuals with head and neck cancer, stratified by ctDNA quantification. [Figure 8]
[0021] This figure shows the age of individuals with renal cell carcinoma (RCC) at the time of sample collection. [Figure 9]
[0022] This figure shows Kaplan-Meier plots illustrating the event-free survival period over time in individuals with RCC at all stages. [Figure 10]
[0023] Figure showing the Kaplan-Meier plot of the event-free survival period over time in individuals with RCC at stages I to III. [Figure 11]
[0024] Figure showing the Kaplan-Meier plot of the recurrence-free survival period over time in individuals with early-stage non-small cell lung cancer (NSCLC). [Figure 12-1]
[0025] Figure showing a schematic diagram of patient treatment, blood collection, sample processing, and the genome-wide methylation enrichment platform. [Figure 12-2]
[0025] Figure showing a schematic diagram of patient treatment, blood collection, sample processing, and the genome-wide methylation enrichment platform. [Figure 13]
[0026] Figure 13A shows a Kaplan-Meier plot indicating the probability of recurrence-free survival over time in individuals predicted to be positive or negative for head and neck cancer based on ctDNA. Figure 13B shows a Kaplan-Meier plot indicating the probability of recurrence-free survival over time in individuals predicted to be positive or negative for head and neck cancer based on ctDNA. Figure 13A shows a Kaplan-Meier plot indicating the probability of recurrence-free survival over time in individuals predicted to be positive or negative for head and neck cancer based on ctDNA at the landmark time point. Figure 13B shows a Kaplan-Meier plot indicating the probability of recurrence-free survival over time in individuals predicted to be positive or negative for head and neck cancer based on ctDNA longitudinally. [Figure 14]
[0027] Figure showing the estimated ctDNA quantification from before treatment to after treatment in head and neck cancer patients with or without recurrence. [Figure 15]
[0028] Figure showing a representative case study of ctDNA dynamics in individual patients before and after treatment for curative purposes. [Figure 16-1]
[0029] Figure showing an example of a schematic diagram of the whole-genome methylation enrichment platform. [Figure 16-2]
[0029] This figure shows an example of a schematic diagram of a whole genome methylation enrichment platform. [Figure 17]
[0030] This figure shows an experimental design for evaluating the detection limits of a whole-genome methylation enrichment platform. [Figure 18]
[0031] This figure shows the methylation specificity distribution and intrinsic molecular distribution of all non-cancer and induced cancer samples treated with the whole-genome methylation enrichment platform disclosed herein. [Figure 19]
[0032] This figure shows the ctDNA methylation scores of different cfDNA sources titrated on pooled non-cancer donor-derived cfDNA in a titration series targeting ctDNA levels of less than 1%. [Figure 20]
[0033] This figure shows an experimental design for evaluating the accuracy of the whole-genome methylation enrichment platform disclosed herein. [Figure 21]
[0034] This figure shows the agreement and variability using ctDNA methylation scores obtained for various levels of ctDNA subjected to the whole-genome methylation enrichment platform disclosed herein. [Figure 22]
[0035] This figure shows the percentage of dispersed components (CV%) for various operators, sequencing runs, and antibody reagent lots, calculated for various levels of ctDNA subjected to the whole-genome methylation enrichment platform disclosed herein. [Figure 23]
[0036] This figure shows the genomic contamination of cell-free DNA. [Figure 24]
[0037] This figure shows the CpG number of 0 for low-binding-specificity and high-binding-specificity samples. [Figure 25-1]
[0038] This figure shows four different workflows for processing plasma-derived cell-free DNA using a whole methylome enrichment platform. [Figure 25-2]
[0038] This figure shows four different workflows for processing plasma-derived cell-free DNA using a total methylome enrichment platform. [Figure 26]
[0039] Figure 25 shows the methylation specificity for different workflows outlined in the diagram. [Figure 27]
[0040] This figure shows the improvement in the detection limit of the cancer methylome approach compared to the whole methylome approach, expressed as a ratio of change. [Figure 28-1]
[0041] This diagram shows a workflow for processing plasma-derived cell-free DNA using a whole methylome enrichment platform, followed by capturing target regions with a probe. [Figure 28-2]
[0041] This figure shows a workflow for processing plasma-derived cell-free DNA using a whole methylome enrichment platform, and then capturing the target region with a probe. [Modes for carrying out the invention]
[0015]
[0042] This disclosure provides methods and systems for processing and analyzing nucleic acids present in biological samples, which may be useful in determining the risk or likelihood of a subject having cancer or tumors with high sensitivity and / or specificity. Methods and systems provided herein may include, for example, the preparation of highly methylated cell-free nucleic acid molecules that can be processed to distinguish cancerous and non-cancerous tissues in circulating free DNA (cfDNA).
[0016]
[0043] For example, the use and analysis of highly methylated nucleic acids can enable highly sensitive and specific detection and / or characterization of circulating tumor DNA (ctDNA) in fluid samples obtained from subjects (e.g., blood samples). In some cases, the use and analysis of highly methylated nucleic acids can enable improved sensitivity, specificity, and / or efficiency in determining the risk of subjects who have or are at risk of developing tumors or cancer. Methods for detecting ctDNA with increased sensitivity are needed, particularly in subjects where ctDNA is present in small amounts.
[0017] definition
[0044] As used herein, the term “subject” generally refers to any component of the animal kingdom (e.g., humans, non-human primates, mice, rats, cattle, sheep, horses, dogs, cats, rabbits, or goats). Therefore, the methods described herein are applicable to both human and veterinary diseases and animal models. A preferred subject is a “patient,” i.e., a living human being being investigated to determine whether treatment or medical care is needed for a disease or condition, or who is receiving medical care for a disease or condition (e.g., cancer).
[0018]
[0045] As used herein, the term “genome” generally refers to genomic information from an object, which may be, for example, at least part or all of the object’s genetic information. A genome can be encoded in either DNA or RNA. A genome may include coding regions (e.g., protein-coding regions) and non-coding regions. A genome can contain the sequences of all chromosomes in an organism. For example, the human genome typically has a total of 46 chromosomes. All of these sequences together can constitute the human genome.
[0019]
[0046] As used herein, the term “nucleic acid” generally refers to a polymeric form of nucleotides of any length, which is a polynucleotide containing two or more nucleotides, i.e., either a deoxyribonucleotide (dNTP) or a ribonucleotide (rNTP), or an analog thereof. Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coding or non-coding regions of genes or gene fragments, loci (may be multiple) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure, where present, may occur before or after the assembly of the nucleic acid. The nucleotide sequence of a nucleic acid may be interrupted by non-nucleotide components. Nucleic acids can be further modified after polymerization, for example, by conjugation or binding with a reporter agent. A “variant” nucleic acid is a polynucleotide having a nucleotide sequence identical to that of the original nucleic acid, except that at least one nucleotide is modified, for example, by deletion, insertion, or substitution. A variant may have a nucleotide sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the nucleotide sequence of the original nucleic acid.
[0020]
[0047] Cell-free methylated DNA (CMS) is DNA that may be one or more nucleic acid molecules freely circulating in the bloodstream. In some cases, CMS can be methylated in various regions of the DNA. Cell-free methylated DNA can be analyzed by taking a sample, such as a plasma sample. Studies have shown that much of the circulating nucleic acid in the blood originates from necrotic or apoptotic cells, and that levels of apoptotic-derived nucleic acid are very high in diseases such as cancer. In particular, for cancer, circulating DNA has prominent signs of the disease, including mutations in oncogenes and microsatellite changes, and for certain cancers, viral genome sequences, DNA, or RNA in plasma are increasingly being studied as potential biomarkers of the disease. For example, quantitative assays for low levels of circulating tumor DNA in total circulating DNA may serve as a better marker for detecting recurrence of colorectal cancer compared to carcinoembryonic antigen, a standard biomarker used clinically. Cell-free DNA (e.g., circulating cfDNA) may include circulating tumor DNA (ctDNA).
[0021]
[0048] As used herein, “library preparation” generally includes one or more of any other preparations performed on cell-free DNA to enable end repair, A-tailing, adapter ligation, or subsequent DNA sequencing.
[0022]
[0049] As used herein, “processed supplemental DNA” (e.g., “filler DNA”) may be non-coding DNA or may consist of amplicons.
[0023]
[0050] In some embodiments, the criterion for measuring fragment length is fragment length. In some preferred embodiments, the cell-free methylated DNA of the present invention is limited to fragments having lengths of less than 170 bp, less than 165 bp, less than 160 bp, less than 155 bp, less than 150 bp, less than 145 bp, less than 140 bp, less than 135 bp, less than 130 bp, less than 125 bp, less than 120 bp, less than 115 bp, less than 110 bp, less than 105 bp, or less than 100 bp. In other preferred embodiments, the cell-free methylated DNA of the present invention is limited to fragments having lengths between about 100 and about 150 bp, 110 and 140 bp, or between 120 and 130 bp.
[0024]
[0051] In some embodiments, the criterion for measuring fragment length is the fragment length distribution of the cell-free methylated DNA in question. In some preferred embodiments, the cell-free methylated DNA in question is limited to fragments within the lower 50th, 45th, 40th, 35th, 30th, 25th, 20th, 15th, or 10th percentiles based on length.
[0025]
[0052] Whenever the terms “at least,” “greater than,” or “greater than or equal to” precede a first number in a set of two or more numbers, the terms “at least,” “greater than,” or “greater than or equal to” apply to each number in that set. For example, 1, 2, or 3 or more is equivalent to 1 or more, 2 or more, or 3 or more.
[0026]
[0053] Whenever the terms “slightly,” “less than,” or “less than or equal to” precede the first number in a set of two or more numbers, the terms “slightly,” “less than,” or “less than or equal to” apply to each of the numbers in that set. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.
[0027] DNA methylation
[0054] Cell-free DNA (cfDNA) can be processed and methylated cfDNA enriched in the cell-free DNA (cfDNA) using any one of the methods disclosed herein. As shown in Figure 1, the cfDNA may undergo end repair, A-tailing, adapter ligation, or other preparations to produce a library of DNA for downstream sequencing. The library may be combined with processed supplemental DNA (e.g., filler DNA) and / or spike-in DNA to produce a sample mixture before thermal denaturation and immunoprecipitation for enriching the methylated cfDNA. Immunoprecipitation may involve combining the sample mixture with any one of the binders disclosed herein and a solid substrate (e.g., multiple magnetic beads). In some cases, immunoprecipitation may involve combining the sample mixture with a methylated nucleic acid capture reagent, which comprises an incubated mixture of the binder and the solid substrate. Upon enrichment of methylated cfDNA, the methylated cfDNA may be amplified and then subjected to sequencing (e.g., an Illumina sequencing reaction).
[0028]
[0055] Cell-free DNA (cfDNA) that may be present in non-invasively collected biological samples (e.g., blood, urine, saliva, cerebrospinal fluid (CSF), etc.) may be a heterogeneous population containing both cfDNA from healthy tissues and cfDNA from tumor or cancer cells (e.g., circulating tumor DNA (ctDNA)). Cancer development may be associated with, for example, a localized increase in 5'-methylcytosine (5mC) in cytosine-phosphate-guanine (CpG) islands and CpG island shores. Cancer development may also be associated with overall (e.g., genome-wide) cytosine demethylation (e.g., overall loss of 5mC). In some cases, ctDNA can be distinguished from cfDNA molecules from healthy tissues (e.g., non-tumor and / or non-cancerous tissues) by the methylation level of the nucleic acid molecule (e.g., the percentage of methylated nucleotide residues). In some cases, nucleic acid molecules from tumor tissue and / or cancerous tissue, or nucleic acid molecules derived therefrom, may be less methylated (e.g., may include lower levels of methylation, e.g., fewer methylated nucleotide residues, and / or a lower percentage of methylated nucleotide residues) compared to nucleic acid molecules from healthy tissue or derived therefrom (e.g., nucleic acid molecules from healthy tissue or derived therefrom consisting of or containing nucleotide sequences corresponding to the same region of the genome in question). For example, a tumor-derived nucleic acid molecule (e.g., a ctDNA molecule) may contain one or more regions with fewer methylated nucleotide residues compared to a nucleic acid molecule (e.g., a cfDNA molecule) derived from healthy tissue (e.g., non-tumor tissue and / or non-cancer tissue) in the same biological sample.In some cases, nucleic acid molecules from tumor tissue and / or cancerous tissue, or nucleic acid molecules derived therefrom, may be more highly methylated (e.g., may include higher levels of methylation, e.g., a greater number of methylated nucleotide residues, and / or a higher percentage of methylated nucleotide residues) compared to nucleic acid molecules from healthy tissue or derived therefrom (e.g., nucleic acid molecules from healthy tissue or derived therefrom consisting of or containing nucleotide sequences corresponding to the same region of the genome in question). For example, a tumor-derived nucleic acid molecule (e.g., a ctDNA molecule) may contain one or more regions with a greater number of methylated nucleotide residues compared to a nucleic acid molecule (e.g., a cfDNA molecule) from healthy tissue (e.g., a non-tumor tissue and / or non-cancer tissue) in the same biological sample. In some cases, all or part of a tumor-derived fraction of multiple cell-free DNA molecules (e.g., ctDNA) can be distinguished from cfDNA molecules derived from healthy tissue by one or more biophysical characteristics (e.g., length of the cfDNA molecule or the presence of typical 5' and 3' terminal sequence motifs) and / or one or more fragmentomics patterns. For example, ctDNA molecules may have shorter nucleic acid lengths than cfDNA molecules derived from healthy tissue. In some cases, ctDNA molecules may contain typical 5' and 3' terminal motifs. In some cases, one or more of these distinctive features can be used to deplete cfDNA from healthy tissue and / or enrich ctDNA in a population of nucleic acid molecules. ctDNA is typically shorter in fragment length compared to cfDNA derived from healthy tissue.
[0029]
[0056] Nucleic acid molecules (e.g., ctDNA) derived from tumor or cancer cells or tissues may be present in biological samples (and / or populations of nucleic acids derived from biological samples) in substantially lower amounts than nucleic acid molecules (e.g., cfDNA) derived from healthy tissue. For example, because the amount of ctDNA present in a sample is lower compared to cfDNA derived from healthy tissue, it may be difficult to detect or sequence (e.g., determine its sequence identity) ctDNA present in or from multiple nucleic acid molecules (e.g., cfDNA) in a biological sample (e.g., it may require the use of larger, potentially rarer biological samples and / or significantly higher sequencing depths).
[0030]
[0057] By depleting (e.g., removing) all or part of a population of methylated DNA molecules (e.g., molecules in which nucleotide methylation levels are increased in all or part of a region of the genome represented by multiple nucleic acid molecules in a biological sample) from multiple nucleic acid molecules (e.g., multiple cell-free nucleic acid molecules or their amplicons containing the biological sample), a remaining population of multiple nucleic acids in the biological sample may be obtained that may be useful in determining the presence and / or sequence identity of ctDNA molecules in the biological sample. Typically, depletion / removal can be carried out by pulling them down using a binder specific to methylated DNA molecules. The pull-down is typically recovered, and the flow-through containing unmethylated / lowmethylated DNA molecules is discarded. This disclosure provides for the first time a method and system for recovering such a flow-through containing unmethylated / lowmethylated DNA molecules and generating a sequencing library using methylated / lowmethylated DNA molecules or derivatives thereof.
[0031]
[0058] In some cases, the depleted sequencing libraries of the methods, systems, compositions, and kits disclosed herein may consist of or comprise such remaining populations of nucleic acid molecules. In some cases, simply depleting nucleic acid molecules methylated in one or more specific regions of the genomic sequence of a nucleic acid molecule (e.g., CpG islands, CpG island shores, or repetitive sequences of the genome, e.g., long scattered repetitive sequences (LINEs), short scattered repetitive sequences (SINEs), or LTRs (terminal repetitive sequences)) from multiple nucleic acids (e.g., cfDNA molecules or their amplicons derived from a biological sample) may be sufficient to achieve increased sensitivity and / or increased specificity in assays for determining the presence or absence or sequence identity of multiple ctDNA molecules. In some cases, multiple nucleic acids (e.g., cfDNA molecules or their amplicons derived from a biological sample) can be subjected to genome-wide depletion of nucleic acid molecules methylated in one or more specific regions of the nucleic acid molecule's genomic sequence (e.g., CpG islands, CpG island shores, or genomic repetitive sequences, e.g., long scattered repetitive sequences (LINEs), short scattered repetitive sequences (SINEs), or LTRs (terminal repetitive sequences)) to potentially achieve increased sensitivity and / or specificity in assays for determining the presence or absence or sequence identity of multiple ctDNA molecules. In some cases, it may be possible to remove CpG genomic islands from the remaining population (e.g., multiple nucleic acid fragments useful in constructing a depleted library). In some cases, the remaining population (e.g., multiple nucleic acid fragments useful in constructing a depleted library) may contain one or more long scattered repetitive sequences (LINEs), short scattered repetitive sequences (SINEs), or terminal repetitive sequences (LTRs) elements.
[0032]
[0059] By enriching all or part of a population of methylated DNA molecules (e.g., molecules in which the level of nucleotide methylation is increased in all or part of a region of the genome represented by multiple nucleic acid molecules in a biological sample) from multiple nucleic acid molecules (e.g., multiple cell-free nucleic acid molecules or their amplicons containing the biological sample), a population of multiple nucleic acids from the biological sample may be obtained that may be useful in determining the presence and / or sequence identity of ctDNA molecules in the biological sample. The enrichment may be carried out by pulling down the methylated DNA molecules using a binder specific to methylated DNA molecules. The pull-down can be recovered, and the flow-through containing unmethylated / low-methylated DNA molecules can be discarded or recovered alternatively (e.g., used to generate a depleted library as described in this disclosure). The enriched fraction can then be subjected to sequencing.
[0033]
[0060] Depletion or enrichment of all or part of methylated nucleic acid molecules among multiple nucleic acid molecules in a biological sample may involve contacting the methylated nucleic acid molecules with a binder (e.g., an affinity molecule such as an antibody or protein specific to methylated nucleotide residues). For example, the preparation of a sequencing library may involve contacting multiple nucleic acid molecules (e.g., cfDNA molecules) or their amplicons with a binder selective to the methylated region of the nucleic acid molecules (e.g., a methylcytosine binder (MBD), e.g., an MBD-Fc fusion protein). In some cases, the binder may be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC)). Cell-free methylated DNA immunoprecipitation sequencing (cfMeDIP-seq), a genome-wide molecular profiling technique, can enrich methylated cfDNA fragments through the use of binders such as anti-5-methylcytosine (anti-5mC) antibodies or methyl-CpG binding domain (MBD) proteins (e.g., MBD-Fc fusion proteins). As described herein, cfMeDIP-seq may include a method and system for depleting methylated DNA fragments from a cfDNA sample, leaving behind hypomethylated or unmethylated cfDNA fragments such as ctDNA. Therefore, the identification of hypomethylated or unmethylated cell-free DNA in a clinical sample may be useful in determining the presence of tumors or cancer in a subject.
[0034]
[0061] In some cases, depletion of multiple nucleic acid molecules (e.g., in the preparation of a depleted sequencing library and / or in the determination of the presence or sequence identity of nucleic acid molecules) may involve removing one or more nucleic acid molecules having methylation levels above a threshold level (e.g., one or more removed nucleic acid molecules are more highly methylated compared to one or more nucleic acid molecules that are not removed during depletion). In some cases, enrichment of multiple nucleic acid molecules (e.g., in the preparation of a enriched sequencing library and / or in the determination of the presence or sequence identity of nucleic acid molecules) may involve removing one or more nucleic acid molecules having methylation levels below a threshold level (e.g., one or more removed nucleic acid molecules are less methylated compared to one or more nucleic acid molecules that are not enriched or methylated). In some cases, the methylation level of a particular nucleic acid fragment (e.g., a DNA fragment) can be considered to have reached a threshold level if a binder having sufficient specificity for methylated cytosine can bind to the particular nucleic acid fragment with or without using filler DNA as described herein. In some cases, the methylation level of a particular nucleic acid fragment (e.g., a DNA fragment) can be considered below a threshold methylation level if a binder with sufficient specificity for methylated cytosine cannot bind to the particular nucleic acid fragment with or without using filler DNA as described herein. In some cases, the depletion of multiple nucleic acid molecules (e.g., in the preparation of a depleted sequencing library and / or the presence or sequencing of nucleic acid molecules) results in (e.g., providing) a remaining population of multiple nucleic acid molecules, which include (or, in some cases, consist of) nucleic acid molecules with methylation levels below a threshold methylation level (e.g., the remaining population is less methylated / less methylated compared to one or more nucleic acid molecules removed from the multiple nucleic acid molecules during depletion). The methylation level can be calculated as the percentage of highly methylated nucleic acid fragments compared to all nucleic acid fragments present in the sample.In some cases, the methylation level thresholds are 0.1%~1%, 1%~5%, 5%~10%, 10%~15%, 15%~20%, 20%~25%, 25%~30%, 30%~35%, 35%~40%, 40%~45%, 45%~50%, 50%~55%, 55%~60%, 65%~70%, 70%~75%, 75%~80%, 80%~85%, 85%~90%, 95%~100%, at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, and less It may be at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, up to 1%, up to 5%, up to 10%, up to 15%, up to 20%, up to 25%, up to 30%, up to 35%, up to 40%, up to 45%, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 95%, or up to 100%.
[0035]
[0062] In some cases, as shown in Figure 1, for example, a first set of nucleic acid molecules (including, for example, nucleic acid molecules such as cfDNA derived from the biological sample in question) may be combined (e.g., mixed) with a second set of nucleic acid molecules (e.g., the second set of nucleic acid molecules are not derived from the subject from which the biological sample was taken). In some cases, the second set of nucleic acid molecules may include processed supplemental DNA (e.g., including lDNA). In some cases, each of the second set of nucleic acid molecules may not be aligned with the human genome.
[0036]
[0063] In some cases, the methods or systems disclosed herein may include, for example, determining or identifying the sequences of all or part of a depleted population of nucleic acid molecules (e.g., the remaining population of nucleic acid fragments of a biological sample after pulling down highly methylated nucleic acid fragments) using a sequencing apparatus. In some cases, the remaining population of nucleic acid molecules may be purified, for example, before or as part of the process of determining or identifying the sequences of all or part of the depleted population of nucleic acid molecules, to obtain a plurality of purified nucleic acid molecules (e.g., after library preparation). In some cases, all or part of the plurality of purified nucleic acid molecules may be amplified (e.g., via polymerase chain reaction) before or as part of the process of determining or identifying the sequences of all or part of the depleted population of nucleic acid molecules. In some cases, the population of amplified nucleic acid molecules or derivatives thereof (e.g., including amplicons of all or part of the plurality of purified nucleic acid molecules) may be subjected to sequencing (e.g., for determining and / or identifying the sequences of nucleic acid molecules). In some cases, sequencing may be achieved using a sequencing apparatus as described herein. In some cases, the sequences of multiple nucleic acid molecules (or their derivatives) in a biological sample can be identified or determined using arrays or polymerase chain reactions. In some cases, the presence of tumor-derived nucleic acid molecules can be determined by calculating the total number of reads per kilobase per million (RPKM) for a region of the genome (e.g., all or part of the genome, e.g., only CpG islands or only CpG island shores). In some cases, the presence of tumor-derived nucleic acid molecules may be indicated if the total RPKM of a depleted sequencing library (e.g., including the remaining population of nucleic acids) is observed to be low, for example, less than 70,000, less than 60,000, less than 50,000, less than 40,000, or less than 30,000 across one or more regions of interest (e.g., CpG islands or CpG island shores).
[0037] Processed supplemental DNA (filler DNA)
[0064] In some cases, processed supplemental DNA (e.g., filler DNA, filler nucleic acid) may be added to a first plurality of nucleic acids (e.g., a plurality of nucleic acids derived from a biological sample, which may include cfDNA from healthy tissue and / or cfDNA from tumor tissue, such as ctDNA). In some cases, the addition of processed supplemental DNA (e.g., a second plurality of nucleic acid molecules) to the first plurality of nucleic acid molecules can increase the specificity and / or sensitivity of the methods, systems, or kits described herein with respect to the detection and / or identification of nucleic acid sequences of the first plurality of nucleic acid molecules, for example. In some cases, the addition of processed supplemental DNA (e.g., a second plurality of nucleic acid molecules) to the first plurality of nucleic acid molecules can increase the rate of methylation region depletion of nucleic acid sequences, for example, during the implementation of some embodiments of the methods and systems described herein. In some cases, the addition of processed supplemental DNA (e.g., a second set of nucleic acid molecules) to a first set of nucleic acid molecules (e.g., including cfDNA of a biological sample) can increase the selectivity of the binder to one or more (e.g., multiple) methylation regions of the first set of nucleic acid molecules. In some cases, the processed supplemental DNA (e.g., a second set of nucleic acid molecules) can be added to the first set of nucleic acid molecules in an amount sufficient to bring the combined mixture of nucleic acid molecules to a desired total mass. In some cases, the desired total mass for use in the methods or systems described herein may be 20ng-30ng, 30ng-40ng, 40ng-50ng, 50ng-60ng, 60ng-70ng, 70ng-80ng, 80ng-90ng, 90ng-100ng, 100ng-110ng, 110ng-120ng, 120ng-130ng, 130ng-140ng, 140ng-150ng, 150ng-160ng, 160ng-170ng, 170ng-180ng, 180ng-190ng, 190ng-200ng, greater than 200ng, or less than 20ng.Depending on the case, 1ng-5ng, 5ng-10ng, 10ng-20ng, 20ng-30ng, 30ng-40ng, 40ng-50ng, 50ng-60ng, 60ng-70ng, 70ng-80ng, 80ng-90ng, 90ng-100ng, 100ng-110ng, 110ng-120ng, 120ng-130ng, 130ng-140ng, 140ng-150ng, 150ng Processed supplemental DNA in amounts of g-160ng, 160ng-170ng, 170ng-180ng, 180ng-190ng, 190ng-200ng, greater than 200ng, less than 20ng, less than 10ng, or less than 5ng can be added to a first plurality of nucleic acid molecules (for example, to bring the total mixture of the processed supplemental DNA and the first plurality of nucleic acid molecules to a desired total mass). In some embodiments, the disclosure includes methods and systems for producing a mixture sample by supplementing a sample with a certain amount of processed supplemental DNA (e.g., filler DNA), the mixture sample containing at least about 50ng, 55ng, 60ng, 65ng, 70ng, 75ng, 80ng, 85ng, 90ng, 95ng, 100ng, 120ng, 140ng, 160ng, 180ng, 200ng, or any amount between a number of the total masses of the nucleic acid mixture. In some embodiments, the processed supplemental DNA comprises at least about 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% methylated processed supplemental DNA, with the remainder being unmethylated processed supplemental DNA, and possibly 5%–50%, 10%–40%, or 15%–30% methylated processed supplemental DNA. In some embodiments, the mixed sample comprises 20ng–100ng, possibly 30ng–100ng, or possibly 50ng–100ng of processed supplemental DNA. In some embodiments, the cell-free DNA derived from the sample and the first amount of processed supplemental DNA together comprise at least 50ng of total DNA, and possibly at least 100ng of total DNA.
[0038]
[0065] In some cases, the processed supplemental DNA may be generated by fragmentation (e.g., via sonication). In some embodiments, the processed supplemental DNA may be 50 bp to 800 bp long, possibly 100 bp to 600 bp long, and possibly 200 bp to 600 bp long. In some embodiments, the processed supplemental DNA is double-stranded. The processed supplemental DNA may be double-stranded DNA. For example, the processed supplemental DNA may be junk DNA. The processed supplemental DNA may also be endogenous or exogenous DNA. For example, the processed supplemental DNA may be non-human DNA, possibly λDNA. As used herein, "λDNA" generally refers to enterobacterial phage λDNA. In some embodiments, the processed supplemental DNA does not substantially sequence-match to human DNA.
[0039]
[0066] In some cases, the processed supplemental DNA (e.g., filler DNA) enriches one or more methylated regions by an enrichment ratio of at least about 1x, at least about 2x, at least about 3x, at least about 4x, at least about 5x, at least about 6x, at least about 7x, at least about 8x, at least about 9x, at least about 10x, at least about 15x, at least about 20x, at least about 25x, at least about 30x, at least about 35x, at least about 40x, at least about 45x, and less Increase by at least 50 times, at least 55 times, at least 60 times, at least 65 times, at least 70 times, at least 75 times, at least 80 times, at least 85 times, at least 90 times, at least 95 times, at least 100 times, at least 150 times, at least 200 times, at least 300 times, at least 400 times, at least 500 times, at least 600 times, at least 700 times, at least 800 times, at least 900 times, or at least 1000 times. In some cases, the processed supplemental DNA (e.g., filler DNA) can be enriched to a maximum of approximately 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 45x, and 50x. It increases by up to approximately 55 times, up to approximately 60 times, up to approximately 65 times, up to approximately 70 times, up to approximately 75 times, up to approximately 80 times, up to approximately 85 times, up to approximately 90 times, up to approximately 95 times, up to approximately 100 times, up to approximately 150 times, up to approximately 200 times, up to approximately 300 times, up to approximately 400 times, up to approximately 500 times, up to approximately 600 times, up to approximately 700 times, up to approximately 800 times, up to approximately 900 times, or up to 1000 times.In some cases, processed supplemental DNA (e.g., filler DNA) can increase the enrichment ratio by approximately 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 55x, 60x, 65x, 70x, 75x, 80x, 85x, 90x, 95x, 100x, 150x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 800x, 900x, or 1000x.
[0040] sample
[0067] The sample may be any biological sample isolated from the subject. For example, the sample may include, but is not limited to, body fluids, whole blood, platelets, serum, plasma, stool, white blood cells or leukocytes, endothelial cells, tissue biopsy, synovial fluid, lymph, ascites, interstitial or extracellular fluid, intercellular fluid including gingival crevicular exudate, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine, fluid from nasal brushing, fluid from papanicolose smear, or any other body fluid. Body fluids may include saliva, blood, or serum. The sample may also be a tumor sample obtained from the subject by a variety of approaches, including but not limited to venous puncture, drainage, ejaculation, massage, biopsy, needle aspiration, irrigation, scraping, surgical incision, or intervention or other approaches. The sample may be a cell-free sample (e.g., substantially free of cells). DNA samples may be denatured, for example, by using sufficient heat.
[0041]
[0068] Samples may be taken from subjects with a disease or disorder. Samples may be taken from subjects suspected of having a disease or disorder. Samples may be taken from subjects with two or more diseases or disorders. Samples may be taken from subjects suspected of having two or more diseases or disorders. In some embodiments, samples may be obtained before and / or after treatment of subjects with a disease or disorder. Samples may be obtained from subjects during treatment or treatment planning. Multiple samples may be obtained from subjects to monitor the effects of treatment over time. The disease or disorder may be cancer.Specific examples of cancer types suitable for detection using the method disclosed herein include acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, basal cell carcinoma, cholangiocarcinoma, bladder cancer, bone cancer, brain tumors, such as cerebellar astrocytoma, cerebral astrocytoma / gliomas, ependymoma, medulloblastoma, supratentorial primordial neuroectodermal tumor, visual pathway and Hypothalamic glioma, breast cancer, colorectal cancer, hepatobiliary tract cancer, bronchial adenoma, Burkitt lymphoma, cancer of unknown primary origin, central nervous system lymphoma, cerebellar astrocytoma, cervical cancer, childhood cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, cutaneous T-cell lymphoma, fibrinogenic round cell tumor, endometrial cancer, ependymoma, esophageal cancer, Ewing's sarcoma, germ cell tumor, gallbladder cancer, gastric cancer Cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, gliomas, hairy cell leukemia, head and neck cancer, cardiac cancer, hepatocellular carcinoma (liver) cancer, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell carcinoma, Kaposi's sarcoma, kidney cancer, laryngeal cancer, lip and oral cancer, liposarcoma, liver cancer, lung cancer, e.g., non-small cell lung cancer and small cell lung cancer, lymphoma, leukemia, macroglobulinemia, malignant fibrous histiocytoma / osteosarcoma of bone, medulloblastoma, melanoma, mesothelioma, metastatic squamous cervical cancer of unknown primary origin, oral cancer, multiple endocrine neoplasia syndrome, myelodysplastic syndrome, myeloid leukemia, nasal and paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, oral cancer Cancer, oropharyngeal cancer, osteosarcoma / malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial carcinoma, ovarian germ cell tumor, pancreatic cancer, pancreatic islet cell tumor, paranasal sinus and nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germ cell tumor, pituitary adenoma, pleuroblastoma, plasma cell tumor, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma, transitional cell carcinoma of the renal pelvis and ureter, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma, skin cancer, Merkel cell carcinoma, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, stomach cancer Examples include cancer, T-cell lymphoma, pharyngeal cancer, thymoma, thymic carcinoma, thyroid cancer, gestational trophoblastic tumor, cancer of unknown primary site, urethral cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenström macroglobulinemia, and Wilms' tumor. In one embodiment, the cancer is head and neck squamous cell carcinoma.
[0042]
[0069] In some embodiments, the cancer may be an early-stage cancer. In some embodiments, the early-stage cancer may be a cancer localized to a small area and / or not having spread to a distal area or nearby tissue. In some embodiments, the early-stage cancer may be a stage I cancer. In some embodiments, the cancer may be a stage I, stage II, stage III, or stage IV cancer.
[0043]
[0070] In some embodiments, the cancer may be a low-exudative tumor (e.g., bladder, breast, endometrium, prostate, or kidney). Low-exudative tumors may release low levels of ctDNA. In some cases, low-exudative tumors may have a low ctDNA load. In some cases, low levels of ctDNA may be found in early-stage cancer.
[0044]
[0071] Samples may be collected from healthy individuals. In some cases, samples may be collected longitudinally from the same individual. In some cases, longitudinally acquired samples may be analyzed for the purpose of monitoring the individual's health and detecting health problems (e.g., early-stage cancer) at an early stage. In some embodiments, samples may be collected at home or at a point of care and then transported by postal service, courier, or other means of transport before analysis. For example, a home user may collect a blood spot sample by fingerprick, which may be dried and then transported by postal service before analysis. In some cases, longitudinally acquired samples may be used to monitor responses to stimuli expected to affect health, motor skills, or cognitive abilities. Non-limiting examples include responses to medication, diet, or exercise management.
[0045]
[0072] In some embodiments, this disclosure provides systems, methods, or kits that include or use one or more biological samples. The one or more samples used herein may include any substance that contains or is presumed to contain nucleic acids. The samples may include biological samples obtained from a subject. In some embodiments, the biological sample is a liquid sample.
[0046]
[0073] In some embodiments, the sample contains an amount of cell-free nucleic acid molecules of about 100 ng, 90 ng, 80 ng, 75 ng, 70 ng, 60 ng, 50 ng, 40 ng, 30 ng, 20 ng, 10 ng, 5 ng, less than 1 ng, or any amount between these values. Furthermore, in some embodiments, the sample contains an amount of cell-free nucleic acid molecules of less than about 1 pg, less than about 5 pg, less than about 10 pg, less than about 20 pg, less than about 30 pg, less than about 40 pg, less than about 50 pg, less than about 100 pg, less than about 200 pg, less than about 500 pg, less than about 1 ng, less than about 5 ng, less than about 10 ng, less than about 20 ng, less than about 30 ng, less than about 40 ng, less than about 50 ng, less than about 100 ng, less than about 200 ng, less than about 500 ng, less than about 1000 ng, or any amount between these values.
[0047]
[0074] In some cases, the preparation or provision of multiple nucleic acid molecules from a biological sample may include performing one or more of the following on the multiple nucleic acid molecules (e.g., after purification from the biological sample): end repair, A-tailing, and adapter ligation.
[0048]
[0075] In some embodiments, a sample may be taken and sequenced at a first time point, and then another sample may be taken and sequenced at a subsequent time point. Such a method may be used, for example, for longitudinal monitoring purposes to track the onset or exacerbation of a disease. In some embodiments, disease exacerbation may be tracked before, after, or during treatment to determine the effectiveness of treatment. For example, the methods described herein may be performed on subjects before and after medical treatment to measure disease exacerbation or regression in response to medical treatment.
[0049]
[0076] After obtaining a sample from a subject, the sample can be processed to generate a dataset that indicates the subject's disease or disorder. For example, the presence, absence, or quantitative evaluation of cell-free nucleic acid molecules (e.g., ctDNA molecules) in the sample in a panel of cancer-related genomic loci or microbiome-related loci may indicate the subject's cancer. The process of processing the sample obtained from the subject may include (i) subjecting the sample to conditions sufficient to isolate, concentrate, or extract multiple cell-free nucleic acid molecules, and (ii) assaying the multiple cell-free nucleic acid molecules to generate a dataset (e.g., nucleic acid sequences). In some embodiments, multiple cell-free nucleic acid molecules are extracted from the sample and subjected to sequencing to generate multiple sequencing reads.
[0050]
[0077] In some embodiments, the cell-free nucleic acid molecule may include cell-free ribonucleic acid (cfRNA) or cell-free deoxyribonucleic acid (cfDNA). The cell-free nucleic acid molecule (e.g., cfRNA or cfDNA) can be extracted from a sample by various methods. The cell-free nucleic acid molecule can be enriched by a group of probes configured to enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to a panel of cancer-associated genomic loci. The probes may have sequence complementarity with nucleic acid sequences from one or more of the panel of cancer-associated genomic loci. A panel of cancer-related genomic loci may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, or more distinct cancer-related genomic loci. Probes may be nucleic acid molecules (e.g., RNA or DNA) that are sequence-complementary to the nucleic acid sequence (e.g., RNA or DNA) of one or more genomic loci (e.g., cancer-related genomic loci). These nucleic acid molecules may be primers or enriched sequences. The step of assaying a sample using a probe that is selective for one or more genomic loci (e.g., cancer-related genomic loci or microbiome-related loci) may include the use of array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing).
[0051]
[0078] Specific methods for capturing cell-free methylated DNA are described in WO2017 / 190215 and WO2019 / 010564, both of which are incorporated by reference in their entirety and for all purposes.
[0052] Sequencing Library
[0079] The methods of this disclosure may utilize various assays, including library preparation (which may include polymerase chain reaction (PCR)) and subsequent sequencing (e.g., next-generation sequencing, Sanger sequencing, etc.). Next-generation sequencing (NGS) technology, also known as high-throughput sequencing, may include various sequencing techniques, including Illumina (Solexa) sequencing, Roche 454 sequencing, Ion torrent:Proton / PGM sequencing, SOLiD sequencing, or long-read sequencing. NGS enables sequencing of DNA and RNA much faster and less expensive than previously used Sanger sequencing. In some embodiments, the sequencing is optimized for short-read sequencing.
[0053]
[0080] Highly methylated sequencing libraries can improve the specificity, sensitivity, and / or efficiency of methods and systems for processing nucleic acids. For example, highly methylated sequencing libraries can improve the specificity, sensitivity, and / or efficiency of assays for determining the presence and / or sequence identity of nucleic acid sequences. A highly methylated sequencing library may contain multiple nucleic acids and / or fragments thereof. In some cases, a highly methylated sequencing library may contain multiple nucleic acid molecules (e.g., a collection of nucleic acids and / or fragments thereof). Multiple nucleic acid molecules may contain all or some of the first multiple nucleic acid molecules, for example, the first multiple nucleic acid molecules may contain one or more nucleic acid molecules containing methylated nucleic acid residues and one or more nucleic acid molecules not containing methylated nucleic acid residues. In some cases, a methylated nucleic acid may contain one or more methylated nucleic acid residues. For example, a methylated nucleic acid may contain one or more methylated cytosines (e.g., one or more 5-methylcytosine (5mC) and / or one or more 5-hydroxymethylcytosine (5hmC)). Multiple nucleic acid molecules (e.g., multiple nucleic acid molecules derived from a biological sample) can be hypermethylated and enriched by using a binder, as described herein, to form a hypermethylated sequencing library that can be used as a novel background, in contrast to a whole-genome background for use in cfDNA analysis. In some cases, DNA may be hypermethylated before the use of a binder in order to construct a sequencing library using a novel background. The novel background sequencing library may contain a set of background genomic regions enriched by the binder.
[0054]
[0081] In some cases, multiple nucleic acids may be subjected to target-specific enrichment. For example, a nucleic acid may be pulled down via a capture probe to enrich a given sequence in a sequencing library. In another example, a nucleic acid may be amplified with a primer to enrich a given sequence in a sequencing library. The capture probe or primer may be specific to a particular gene, non-coding region, or other sequence. For example, one or more genes may be enriched in a nucleic acid. One or more genes may be cancer-related genes. For example, one or more genes may be genes previously determined to be associated with a type of cancer. For example, a gene or sequence having a previously identified methylation state or a previously identified number of methylated nucleotides may be enriched in a nucleic acid. Target-specific enrichment may be performed during the preparation of the sequencing library. For example, a nucleic acid may be enriched or depleted from methylated nucleic acids (as described in this disclosure), and then the nucleic acid may be subjected to target-specific enrichment.
[0055] Nucleic acid molecule sequencing
[0082] This disclosure provides methods and techniques for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides may be nucleic acid molecules, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), and include variants or derivatives thereof (e.g., single-stranded DNA). Sequencing may be performed by next-generation sequencing.
[0056]
[0083] Furthermore, any sequencing method that provides fragment lengths, such as paired-end sequencing, can be utilized. Alternatively, sequencing may be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such systems may provide multiple raw genetic data corresponding to the genetic information of a subject (e.g., a human) to be generated by the system from a sample provided by the subject. In some examples, such systems provide sequencing reads (also referred to herein as “reads”). Reads may contain strings of nucleic acid bases corresponding to the sequence of the sequenced nucleic acid molecule. In some circumstances, the systems and methods provided herein may be used in conjunction with proteomic information.
[0057]
[0084] In some embodiments, sequencing reads are obtained via next-generation sequencing or next-next-generation sequencing. In some embodiments, the sequencing method includes cfMeDIP sequencing, which includes a process or system such as that described by Shen et al. ("Sensitive tumor detection and classification using plasma cell-free DNA methylomes", (2018) Nature), which is incorporated entirely herein. In some embodiments, sequencing may be whole-genome sequencing. In some embodiments, sequencing may be whole-exome sequencing. In some embodiments, sequencing may be partial-genome sequencing. In some embodiments, sequencing may be target genome sequencing. In some embodiments, sequencing may be whole-methylome sequencing. In some embodiments, sequencing may be whole-methylome sequencing rather than whole-genome sequencing. In some embodiments, whole-methylome sequencing can be captured without requiring a predefined target panel. In some embodiments, sequencing may be sequencing of a partial methylome. In some embodiments, sequencing may be sequencing of a target methylome. In some embodiments, sequencing may be performed using methyl-CpG binding domain sequencing (MBD-seq). In some cases, MBD-seq may include capturing double-stranded methylated DNA fragments for sequencing a methylation-enriched DNA fragment library (e.g., via a binder such as an antibody specific to the species of methylated nucleotide). In some embodiments, the sequencing method includes Cancer Personalized Profiling by deep Sequencing (CAPP-Seq), a next-generation sequencing-based method used to quantify circulating DNA (ctDNA) in cancer.This method can be generalized to any type of cancer recorded as having recurrent mutations and can detect one mutant DNA molecule in 10,000 healthy DNA molecules. In some embodiments, sequencing may include chemical transformations. In some embodiments, sequencing may include bisulfite sequencing. In some embodiments, sequencing may not include bisulfite sequencing.
[0058]
[0085] Sequencing may include targeted sequencing. For example, a sequencing reaction may include a capture probe specific to the region of interest. The use of targeted sequencing can increase the amount of region-specific reads that are useful (e.g., those related to DMR or usable to distinguish healthy subjects from diseased subjects). A capture probe may include one or more probes complementary or homologous to a region having one or more sites suitable for enzymatic methylation. A capture probe may include one or more probes complementary or homologous to a region having one or more sites that are substantially unmethylated in healthy controls or disease-free controls. A capture probe may include one or more probes complementary or homologous to a region having a known methylation state. Targeted sequencing may target one or more regions known to be present in multiple types of cancer.
[0059]
[0086] In some embodiments, sequencing preserves fragment length and / or terminal motifs. Non-limiting examples of terminal motifs include 5' terminal motifs, 3' terminal motifs, 6-mer terminal motifs, 5-mer terminal motifs, and 4-mer terminal motifs (e.g., CCCA, CCTG, CCAG, CCAA, CCAT, CCTG, CCAA, CCCT, CCTC, TGTG, TGTT, CCTA, TATT, CCAC, TCTT, CCCC, TATA, TAAA, AAAA, TTTT, or other variations thereof). Preservation of fragment length and terminal motifs can enable multi-omics evaluation in a single assay, increase efficiency at low concentrations of ctDNA, and enhance performance. In some embodiments, sequencing may include enzymatic conversion. In some cases, enzymatic conversion may include TET2 and an oxidative enhancer. In some cases, enzymatic conversion may further include APOBEC. In some embodiments, sequencing does not include enzymatic conversion. In some embodiments, the quality of the DNA and the four DNA bases (e.g., adenine, cytosine, guanine, and thymine) can be preserved by not performing any conversion (e.g., chemical or enzymatic). By preserving DNA quality, more reads can pass through a quality control filter that is uniquely aligned to the human genome.
[0060]
[0087] In some cases, the sample or a portion of it (e.g., multiple nucleic acids from the sample) may be subjected to library preparation before sequencing. In short, after end repair and A-tailing, the sample is ligated to a nucleic acid adapter and digested using enzymes.
[0061]
[0088] In some embodiments, sequencing includes modification of a nucleic acid molecule or fragment by ligating it with, for example, a barcode, molecular barcode (UMI), or another tag. Ligating a barcode, UMI, or tag to one end of a nucleic acid molecule or fragment can facilitate analysis of the nucleic acid molecule or fragment after sequencing. In some embodiments, the barcode is a unique barcode (e.g., a UMI). In some embodiments, the barcode is non-unique, and the barcode sequence may be used in relation to endogenous sequence information, such as the start and stop sequences of the target nucleic acid (e.g., the target nucleic acid is surrounded on both sides by the barcode and barcode sequence, generating a uniquely tagged molecule in relation to the start and end sequences of the target nucleic acid). The barcode, UMI, or tag may be a known sequence used to associate a polynucleotide or fragment thereof with an input or target nucleic acid molecule or fragment thereof. The barcode, UMI, or tag may include natural nucleotides or non-natural (e.g., modified) nucleotides (e.g., as described herein). The barcode sequence may be included within the adapter sequence so that the barcode sequence can be included within the sequencing read. The barcode sequence may contain at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more nucleotide lengths. In some cases, the barcode sequence may be long enough to enable the identification of a sample based on a related barcode sequence, and may be sufficiently different from another barcode sequence. The barcode sequence, or a combination of barcode sequences, can be used to tag and subsequently identify the “original” nucleic acid molecule or fragment thereof (e.g., a nucleic acid molecule or fragment thereof present in a sample from which the subject originates). In some cases, the barcode sequence, or a combination of barcode sequences, can be used in conjunction with endogenous sequence information to identify the original nucleic acid molecule or fragment thereof. For example, the barcode sequence, or a combination of barcode sequences, may be used with the endogenous sequence adjacent to the barcode, UMI, or tag (e.g., the start and end positions of the endogenous sequence).
[0062]
[0089] As described herein, the prepared library can be combined with filler nucleic acids (e.g., filler λDNA) to minimize the impact of low-abundance ctDNA in the prepared library and generate a mixed sample. In some embodiments, when the disease / condition is a localized (non-metastatic) cancer, the amount of ctDNA may be low and may not be easily and accurately measured and quantified. In such cases, the mixed sample is concentrated to at least about 50 ng, 80 ng, 100 ng, 120 ng, 150 ng, or 200 ng and subjected to further concentration.
[0063]
[0090] The process of processing nucleic acid molecules or fragments may include performing nucleic acid amplification. Amplification of nucleic acids (e.g., enriched methylated nucleic acids) before sequencing may produce more sequence reads. Amplification may produce more stable nucleic acids (e.g., partially single-stranded or double-stranded nucleic acids compared to single-stranded nucleic acids), which can then be stored for longer periods without degradation. For example, any type of nucleic acid amplification reaction can be used to amplify a target nucleic acid molecule or fragment and produce an amplification product. Non-limiting examples of nucleic acid amplification methods include reverse transcription, primer extension, polymerase chain reaction (PCR), ligase chain reaction, asymmetric amplification, rolling circle amplification, and multiple substitution amplification (MDA). Examples of PCR include, but are not limited to, quantitative PCR, real-time PCR, digital PCR, emulsion PCR, hot-start PCR, multiplex PCR, asymmetric PCR, nested PCR, and assembly PCR. Nucleic acid amplification may involve one or more primers, probes, polymerases, buffers, enzymes, and one or more reagents such as deoxyribonucleotides. Nucleic acid amplification may be isothermal, or may involve thermal cycling, and / or may have the length of the endogenous sequence. In some cases, PCR amplification may involve at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification.
[0064]
[0091] Amplification for generating nucleic acids suitable for sequencing may be performed on nucleic acids bound to or attached to a solid substrate (e.g., beads). For example, as described elsewhere in this disclosure, methylation molecules may be attached to the solid substrate, and methylation molecules may bind to nucleic acids. The bound nucleic acids may be amplified to produce amplicons that are not bound to the binding molecules or the solid substrate. Amplification of nucleic acids bound to a solid substrate may enable improved throughput, for example, by reducing the need for washing or buffer exchange that may occur during the elution or removal of nucleic acids from the solid substrate.
[0065]
[0092] Processing of nucleic acid molecules or fragments may involve the addition of beads capable of binding to nucleic acids. Nucleic acid binding may allow for specific binding of the nucleic acid and the removal of enzymes or contaminants from the nucleic acid sample. Beads (e.g., SPRI beads) can bind to nucleic acids in the library and can be washed or separated to remove contaminants. Nucleic acids can be eluted from the beads and then collected. This can be done by removing the beads from the nucleic acid sample. For example, the beads may be magnetic beads and may be subjected to a magnetic field to remove them. The removal process may be repeated multiple times to reduce or minimize the carryover of beads to subsequent processing reactions.
[0066] Binder
[0093] A binder may be used to deplete or enrich a population of nucleic acid molecules (e.g., multiple nucleic acid molecules derived from a biological sample). In some cases, a binder may be used to deplete or enrich multiple nucleic acid molecules of one or more nucleic acid molecules having a methylation level above a threshold level (e.g., by binding to one or more methylated nucleotides of one or more nucleic acid molecules). A binder may be used to enrich a population of nucleic acid molecules (e.g., multiple nucleic acids derived from a biological sample). The binder may be a molecule that specifically binds to methylated nucleic acids or methylated nucleotides. In some cases, the binder may be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 4-methylcytosine (4mC), or 6-methyladenine (6mA)). In some cases, the binder can be selected from the group consisting of anti-5-methylcytosine antibodies or derivatives thereof, anti-5-carboxylcytosine antibodies or derivatives thereof, anti-5-formylcytosine antibodies or derivatives thereof, anti-5-hydroxymethylcytosine antibodies or derivatives thereof, anti-3-methylcytosine antibodies or derivatives thereof, and any combination thereof. In some cases, the binder may be an anti-5-methylcytosine antibody or a derivative thereof. In some embodiments, the binder is a protein containing a methyl-CpG binding domain. One such protein is the MBD2 protein. As used herein, “methyl-CpG binding domain (MBD)” generally refers to a specific domain of proteins and enzymes that binds to DNA containing one or more symmetrically methylated CpGs, and is approximately 70 residues long. The MBDs of MeCP2, MBD1, MBD2, MBD4, and BAZ2 mediate binding to DNA, and in the case of MeCP2, MBD1, and MBD2, preferentially mediate binding to methylated CpGs. The human proteins MECP2, MBD1, MBD2, MBD3, and MBD4 comprise a family of nuclear proteins that are related in that each contains a methyl-CpG binding domain (MBD). Each of these proteins, with the exception of MBD3, can specifically bind to methylated DNA.The binder may contain biotin, which may enable the binder to couple to or bind to streptavidin.
[0067]
[0094] In other embodiments, the binder is an antibody, and the step of capturing cell-free methylated DNA includes immunoprecipitation of the cell-free methylated DNA using the antibody. As used herein, “immunoprecipitation” generally refers to the technique of precipitating an antigen (such as polypeptides and nucleotides) from a solution using an antibody that specifically binds to a particular antigen. This process may also be used to isolate and concentrate a particular protein or DNA from a sample, and may require the antibody to be bound to a solid substrate at some point in this procedure. The solid substrate includes, for example, beads such as magnetic beads. Other types of beads and solid substrates may be used. Various proteins or chemical moieties that can enable coupling or binding to the binder can be present on the solid substrate. For example, the solid substrate can bind to or couple to a molecule that specifically binds to methylated nucleic acids. For example, the solid substrate may contain streptavidin, which can bind to a biotinylated antibody. In another example, the solid substrate may contain protein A, which can bind to the Fc domain of an antibody.
[0068]
[0095] In various embodiments, a binder (e.g., a methylation molecule) is added to the sample to bind to nucleic acids. The binder may be bound to a solid substrate (e.g., magnetic beads), and this complex (e.g., a methylated nucleic acid capture reagent) may be added to the sample. The complex may first be generated by pre-incubation in the absence of the sample. For example, an anti-5-mC antibody may be incubated with protein A beads, allowing the antibody to bind or couple to the protein A beads via its Fc domain. This allows the beads to be saturated with the antibody, and the complex can bind to methylated nucleic acids. Once the complex is generated, it may be added to the sample to bind to nucleic acids. The binder and solid substrate may also be added to the sample simultaneously (or substantially simultaneously), or sequentially. For example, magnetic protein A beads and an antibody may be added to the sample simultaneously, allowing the antibody to bind to the nucleic acid and the antibody to the magnetic protein A beads. In another example, the antibody may be added to the sample first to bind the nucleic acid to the antibody, and then the beads may be added to bind to the antibody bound to the nucleic acid. Each method of complex formation—pre-incubation, simultaneous addition, or sequential addition—may offer its own advantages. For example, pre-incubation allows for the formation of a stable complex between the binder and the solid substrate without the potential steric hindrance of the nucleic acid.
[0069]
[0096] For example, a 5-mC antibody (e.g., a 5-mC antibody that specifically binds to 5-methylcytosine) can be used as a binder. For the immunoprecipitation procedure, in some embodiments, at least 0.05 μg of antibody is added to the sample, while in some embodiments, at least 0.16 μg of antibody is added to the sample. In some cases, 0.05 μg to 0.80 μg, 0.16 μg to 0.80 μg, 0.40 μg to 0.80 μg, 0.16 μg to 0.40 μg, 0.10 μg to 0.80 μg, 0.20 μg to 0.60 μg, 0.30 μg to 0.50 μg, or 0.40 μg to 0.50 μg of antibody can be used. To confirm the immunoprecipitation reaction, in some embodiments, the method described herein further includes the step of adding a second amount of control DNA to the sample.
[0070]
[0097] In some embodiments, the immunoprecipitation process is optimized, and this optimization may include changing the binder used (e.g., antibody), adjusting the concentration of the binder, and / or adjusting the length of time that the binder is allowed to capture the cell-free methylated DNA.
[0071] Processing method for methylated nucleic acids
[0098] This disclosure provides a method and system for processing cell-free nucleic acid samples from a subject to detect methylation events.
[0072]
[0099] In some embodiments, as shown in Figure 25 (Workflow 1), the method includes the steps of: (a) providing a plurality of nucleic acid molecules derived from a nucleic acid sample (e.g., cfDNA, cfDNA having spike-in DNA); (b) generating a library by library preparation (e.g., end repair, A-tail, adapter ligation) using one or more custom adapters; (c) generating a sample mixture by adding a plurality of filler nucleic acid molecules; (d) thermally denaturing and rapidly cooling the sample mixture; (e) subjecting the sample mixture to immunoprecipitation to obtain a concentrated sample containing a plurality of methylated nucleic acid molecules; and (f) preparing the concentrated sample for sequencing. In some cases of this method, (e) the step of subjecting the sample mixture to immunoprecipitation further includes i) adding a binder disclosed herein to the sample mixture, ii) adding a solid substrate (e.g., a magnetic solid substrate) to the sample mixture, iii) incubating the binder and solid substrate with the sample mixture for a sufficient amount of time (e.g., overnight) to capture a concentrated sample containing multiple methylated nucleic acid molecules, and iii) isolating the binder and solid substrate from the concentrated sample. In some cases of this method, (f) the step of preparing a concentrated sample for sequencing further includes i) cleaning up the concentrated sample to produce a cleaned concentrated sample, ii) performing PCR amplification on the cleaned concentrated sample to obtain multiple PCR amplicons, and iii) cleaning up the PCR amplicons before sequencing. In some cases, the PCR amplification may be 14 cycles.
[0073]
[0100] In some embodiments, as shown in Figure 25 (Workflow 2), the method includes (a) providing a plurality of nucleic acid molecules derived from a nucleic acid sample (e.g., cfDNA, cfDNA having spike-in DNA); (b) preparing a library by library preparation (e.g., end repair, A-tail, adapter ligation) using one or more custom adapters; (c) performing post-library preparation cleanup using a solid substrate (e.g., magnetic solid substrate); (d) adding a plurality of filler nucleic acid molecules to produce a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) subjecting the sample mixture to immunoprecipitation to obtain a concentrated sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the concentrated sample for sequencing. In some cases of the method, step (c) performing post-library preparation cleanup includes further capture of the solid substrate (e.g., magnetic solid substrate) by placing the prepared library in a device (e.g., magnetic rack) to capture any remaining solid substrate. In some cases of this method, (f) the step of subjecting the sample mixture to immunoprecipitation further includes i) adding a binder disclosed herein to the sample mixture, ii) adding a solid substrate (e.g., a magnetic solid substrate) to the sample mixture, iii) incubating the binder and solid substrate with the sample mixture for a sufficient amount of time (e.g., overnight) to capture a concentrated sample containing multiple methylated nucleic acid molecules, and iii) isolating the binder and solid substrate from the concentrated sample. In some cases of this method, (g) the step of preparing a concentrated sample for sequencing further includes i) cleaning up the concentrated sample to produce a cleaned concentrated sample, ii) performing PCR amplification on the cleaned concentrated sample to obtain multiple PCR amplicons, and iii) cleaning up the PCR amplicons before sequencing. In some cases, the PCR amplification may be 14 cycles.
[0074]
[0101] In some embodiments, as shown in Figure 25 (Workflow 3), the method includes (a) providing a plurality of nucleic acid molecules derived from a nucleic acid sample (e.g., cfDNA, cfDNA having spike-in DNA); (b) preparing a library by library preparation (e.g., end repair, A-tail, adapter ligation) using one or more custom adapters; (c) performing post-library preparation cleanup using a solid substrate (e.g., magnetic solid substrate); (d) adding a plurality of filler nucleic acid molecules to produce a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) subjecting the sample mixture to immunoprecipitation to obtain a concentrated sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the concentrated sample for sequencing. In some cases of the method, step (c) performing post-library preparation cleanup includes further capture of the solid substrate (e.g., magnetic solid substrate) by placing the prepared library in a device (e.g., magnetic rack) to capture any remaining solid substrate. In some cases of this method, (f) the step of subjecting the sample mixture to immunoprecipitation further includes i) adding a binder disclosed herein to the sample mixture, ii) adding a solid substrate (e.g., a magnetic solid substrate) to the sample mixture, iii) incubating the binder and solid substrate with the sample mixture for a sufficient amount of time (e.g., overnight) to capture a concentrated sample containing multiple methylated nucleic acid molecules, and iii) isolating the binder and solid substrate from the concentrated sample. In some cases of this method, (g) the step of preparing a concentrated sample for sequencing further includes i) cleaning up the concentrated sample to produce a cleaned concentrated sample, ii) performing PCR amplification on the cleaned concentrated sample to obtain multiple PCR amplicons, and iii) cleaning up the PCR amplicons before sequencing. In some cases, the PCR amplification may be 13 cycles.
[0075]
[0102] In some embodiments, as shown in Figure 25 (Workflow 4), the method includes (a) providing a plurality of nucleic acid molecules derived from a nucleic acid sample (e.g., cfDNA, cfDNA having spike-in DNA); (b) preparing a library by library preparation (e.g., end repair, A-tail, adapter ligation) using one or more custom adapters; (c) performing post-library preparation cleanup using a solid substrate (e.g., magnetic solid substrate); (d) adding a plurality of filler nucleic acid molecules to produce a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) subjecting the sample mixture to immunoprecipitation to obtain a concentrated sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the concentrated sample for sequencing. In some cases of the method, step (c) performing post-library preparation cleanup includes further capture of the solid substrate (e.g., magnetic solid substrate) by placing the prepared library in a device (e.g., magnetic rack) to capture any remaining solid substrate. In some cases of this method, (f) the step of subjecting the sample mixture to immunoprecipitation further includes i) incubating the binder disclosed herein with a solid substrate (e.g., a magnetic solid substrate) to produce a methylated nucleic acid capture reagent; ii) adding the methylated nucleic acid capture reagent to the sample mixture; and ii) after incubation the methylated nucleic acid capture reagent and the sample mixture for a sufficient time (e.g., overnight), isolating the methylated nucleic acid capture reagent to obtain a concentrated sample containing multiple methylated nucleic acid molecules. In some cases of this method, (g) the step of preparing a concentrated sample for sequencing further includes i) cleaning up the concentrated sample to produce a cleaned concentrated sample; ii) performing PCR amplification on the cleaned concentrated sample to obtain multiple PCR amplicons; and iii) cleaning up the PCR amplicons before sequencing. In some cases, the PCR amplification may be 13 cycles.
[0076]
[0103] In some embodiments, as shown in Figure 28, the method includes (a) providing a plurality of nucleic acid molecules derived from a nucleic acid sample (e.g., cfDNA, cfDNA having spike-in DNA); (b) preparing a library by library preparation (e.g., end repair, A-tail, adapter ligation) using one or more custom adapters; (c) performing post-library preparation cleanup using a solid substrate (e.g., magnetic solid substrate); (d) adding a plurality of filler nucleic acid molecules to produce a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) subjecting the sample mixture to immunoprecipitation to obtain a concentrated sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the concentrated sample for sequencing. In some cases of the method, step (c) performing post-library preparation cleanup includes further capture of the solid substrate (e.g., magnetic solid substrate) by placing the prepared library in a device (e.g., magnetic rack) to capture any remaining solid substrate. In some cases of this method, (f) the step of subjecting the sample mixture to immunoprecipitation further includes i) incubating a binder disclosed herein with a solid substrate (e.g., a magnetic solid substrate) to produce a methylated nucleic acid capture reagent; ii) adding the methylated nucleic acid capture reagent to the sample mixture; and ii) after incubation the methylated nucleic acid capture reagent and the sample mixture for a sufficient time (e.g., overnight), isolating the methylated nucleic acid capture reagent to obtain a concentrated sample containing multiple methylated nucleic acid molecules. In some cases of this method, (g) the step of preparing a concentrated sample for sequencing further includes i) cleaning up the concentrated sample to produce a cleaned concentrated sample; ii) performing PCR amplification on the cleaned concentrated sample to obtain multiple PCR amplicons; iii) cleaning up the PCR amplicons; and iv) contacting the multiple methylated nucleic acids with one or more nucleic acid capture probes to concentrate one or more target sequences before sequencing. In some cases, the PCR amplification may be 13 cycles. In some cases, one or more target sequences contain one or more genes.
[0077]
[0104] In one embodiment, a method or system for processing a nucleic acid sample from a subject is disclosed herein, comprising the steps of (a) generating a nucleic acid sample mixture containing multiple methylated nucleic acids from the subject and a certain amount of processed supplemental DNA (e.g., filler DNA); (b) incubating the nucleic acid sample mixture with (i) a methylation-binding molecule and (ii) a solid substrate; and (c) capturing the methylated nucleic acids to concentrate multiple methylated nucleic acids (e.g., methylated single-stranded DNA) in the nucleic acid sample mixture. In some cases, a certain amount of processed supplemental DNA (e.g., filler DNA) is not required in (a). In some cases, a certain amount of processed supplemental DNA (e.g., filler DNA) contains at least one methylated DNA molecule. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule contains biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the solid substrate is beads. In some cases, the solid substrate is a magnetic solid substrate. In some cases, the solid substrate contains protein A. In some cases, the solid substrate contains streptavidin. In some cases, the method further includes a step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further includes a step of obtaining a nucleic acid sample from the sample and a step of performing one or more library preparation reactions on the nucleic acids before (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions and before (a), the prepared library is incubated with a plurality of DNA capture beads (e.g., SPRI beads) and then removed. In some cases, the method further includes a step of amplifying the methylated nucleic acids after the capture step to produce amplicons of a plurality of methylated nucleic acids. In some cases, the captured methylated nucleic acids are subjected to buffer exchange or washing reactions before amplification.In some cases, the captured methylated nucleic acid is subjected to an elution reaction before amplification. In some cases, the captured methylated nucleic acid is not subjected to an elution reaction before amplification. In some cases, amplification is via PCR amplification. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is 13 cycles of amplification. In some cases, the amplicon generated from the methylated nucleic acid is subjected to a sequencing reaction. In some cases, the amplicon undergoes cleanup before sequencing. In some cases, the amplicon does not undergo cleanup before sequencing. In some cases, the sequencing reaction is sequencing by a synthetic reaction. In some cases, the sequencing reaction does not include bisulfite sequencing. In some cases, the method or system further includes the step of contacting multiple methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, one or more target sequences contain one or more genes.
[0078]
[0105] In one embodiment, a method or system for processing a nucleic acid sample from a subject is disclosed herein, comprising the steps of (a) generating a nucleic acid sample mixture comprising multiple methylated nucleic acids from a subject and a certain amount of processed supplemental DNA (e.g., filler DNA); (b) (i) incubating a methylation-binding molecule with (ii) a solid substrate to form a methylated nucleic acid capture reagent; and (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture to concentrate multiple methylated nucleic acids in the nucleic acid mixture. In some cases, a certain amount of processed supplemental DNA (e.g., filler DNA) is not required in (a). In some cases, the certain amount of processed supplemental DNA (e.g., filler DNA) comprises at least one methylated DNA molecule. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule comprises biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the solid substrate is beads. In some cases, the solid substrate is a magnetic solid substrate. In some cases, the solid substrate contains protein A. In some cases, the solid substrate contains streptavidin. In some cases, the method further comprises a step of denaturing the nucleic acids in the nucleic acid sample mixture before (b). In some cases, the method further comprises a step of obtaining a nucleic acid sample from the sample and a step of carrying out one or more library preparation reactions on the nucleic acids before (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before carrying out one or more library preparation reactions. In some cases, after carrying out one or more library preparation reactions and before (a), the prepared library is incubated with a number of magnetic beads that interact with nucleic acids (e.g., SPRI beads) and eluted from the magnetic beads that interact with nucleic acids. In some cases, the sample is subjected to magnetic capture to remove the magnetic beads that interact with nucleic acids. In some cases, the sample is subjected to additional magnetic capture to remove residual magnetic beads that interact with nucleic acids.In some cases, the method further includes a step of amplifying the methylated nucleic acid after the capture step to generate amplicons of multiple methylated nucleic acids. In some cases, the amplification step is performed while the methylated nucleic acid is bound to the methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acid is subjected to a buffer exchange or washing reaction before amplification. In some cases, the captured methylated nucleic acid is subjected to an elution reaction before amplification. In some cases, the captured methylated nucleic acid is not subjected to an elution reaction before amplification (for example, if amplification is performed while the methylated nucleic acid is bound to the methylated nucleic acid capture reagent). In some cases, amplification is via PCR amplification. In some cases, the PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, the PCR amplification is 14 cycles of amplification. In some cases, the PCR amplification is 13 cycles of amplification. In some cases, the generated amplicons of methylated nucleic acids are subjected to a sequencing reaction. In some cases, the amplicons undergo cleanup before sequencing. In some cases, the amplicon does not undergo cleanup before sequencing. In some cases, the sequencing reaction is synthetic sequencing. In some cases, the sequencing reaction does not involve bisulfite sequencing. In some cases, the method or system further includes a step of enriching one or more target sequences by contacting multiple methylated nucleic acids with one or more nucleic acid capture probes. In some cases, one or more target sequences include one or more genes.
[0079]
[0106] In one embodiment, a method or system for processing a nucleic acid sample from a subject is disclosed herein, comprising the steps of (a) providing a nucleic acid sample containing a plurality of methylated nucleic acids; (b) (i) incubating methylation-binding molecules with (ii) a solid substrate to form a methylated nucleic acid capture reagent; (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating methylated nucleic acids bound to a solid substrate; and (d) amplifying the methylated nucleic acids bound to the solid substrate to generate an amplicon of the methylated nucleic acids. In some cases, the amplification step is performed while the methylated nucleic acids are bound to the methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some cases, the methylation-binding molecules are antibodies (e.g., anti-5-methylcytosine (anti-5mC) antibodies, methyl-CpG-binding domain (MBD) proteins). In some cases, the methylation-binding molecules include biotin. In some cases, the methylation-binding molecules bind to methylated cytosine. In some cases, the solid substrate is beads. In some cases, the solid substrate is a magnetic solid substrate. In some cases, the solid substrate contains protein A. In some cases, the solid substrate contains streptavidin. In some cases, the method further includes a step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further includes a step of obtaining a nucleic acid sample from the sample and a step of performing one or more library preparation reactions on the nucleic acids before (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions and before (a), the prepared library is incubated with a number of magnetic beads that interact with nucleic acids (e.g., SPRI beads) and eluted from the magnetic beads that interact with nucleic acids. In some cases, the sample is subjected to magnetic capture to remove the magnetic beads that interact with nucleic acids. In some cases, the sample is subjected to additional magnetic capture to remove residual magnetic beads that interact with nucleic acids. In some cases, the captured methylated nucleic acids are not subjected to elution reactions before amplification.In some cases, amplification is carried out via PCR amplification. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is 13 cycles of amplification. In some cases, the amplicons generated from methylated nucleic acids are subjected to a sequencing reaction. In some cases, the amplicons undergo cleanup before sequencing. In some cases, the amplicons do not undergo cleanup before sequencing. In some cases, the sequencing reaction is sequencing by a synthetic reaction. In some cases, the sequencing reaction does not involve bisulfite sequencing. In some cases, the method or system further includes a step of contacting multiple methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, one or more target sequences include one or more genes.
[0080]
[0107] In one embodiment, a method or system for processing a nucleic acid sample from a subject is disclosed herein, comprising the steps of: (a) generating a nucleic acid sample mixture comprising multiple methylated nucleic acids from the subject and a certain amount of processed supplemental DNA (e.g., filler DNA), wherein the filler DNA comprises at least one methylated DNA molecule; (b) capturing the methylated nucleic acids by adding a capture reagent comprising a solid substrate to the nucleic acid sample mixture, thereby generating methylated nucleic acids bound to the solid substrate; and (c) amplifying the methylated nucleic acids bound to the solid substrate to generate an amplicon of the methylated nucleic acid. In some cases, the certain amount of processed supplemental DNA (e.g., filler DNA) comprises at least one methylated DNA molecule. In some cases, the amplification step is performed while the methylated nucleic acids are bound to the methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acids are not subjected to an elution reaction before amplification. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation molecule contains biotin. In some cases, the methylation molecule binds to methylated cytosine. In some cases, the solid substrate is beads. In some cases, the solid substrate is a magnetic solid substrate. In some cases, the solid substrate contains protein A. In some cases, the solid substrate contains streptavidin. In some cases, the method further comprises a step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further comprises a step of obtaining a nucleic acid sample from the sample and a step of performing one or more library preparation reactions on the nucleic acids before (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions and before (a), the prepared library is incubated with a number of magnetic beads that interact with nucleic acids (e.g., SPRI beads) and eluted from the magnetic beads that interact with nucleic acids.In some cases, the sample is subjected to magnetic capture to remove magnetic beads that interact with the nucleic acid. In some cases, the sample is subjected to additional magnetic capture to remove residual magnetic beads that interact with the nucleic acid. In some cases, the captured methylated nucleic acid is not subjected to an elution reaction before amplification. In some cases, amplification is via PCR amplification. In some cases, PCR amplification involves at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is 13 cycles of amplification. In some cases, the amplicon generated from the methylated nucleic acid is subjected to a sequencing reaction. In some cases, the amplicon undergoes cleanup before sequencing. In some cases, the amplicon does not undergo cleanup before sequencing. In some cases, the sequencing reaction is sequencing by a synthetic reaction. In some cases, the sequencing reaction does not involve bisulfite sequencing. In some cases, the method or system further includes a step of enriching one or more target sequences by contacting multiple methylated nucleic acids with one or more nucleic acid capture probes. In some cases, the one or more target sequences include one or more genes.
[0081]
[0108] In one embodiment, a method or system is disclosed herein that includes (a) obtaining a first nucleic acid molecule from a target cell-free sample; (b) generating a second set of nucleic acid molecules from the first set of nucleic acid molecules or derivatives thereof, wherein the methylation level of the second set of nucleic acid molecules is concentrated compared to that of the first set of nucleic acid molecules; (c) concentrating one or more targets in the second set of nucleic acid molecules or derivatives thereof to obtain a third set of nucleic acid molecules; and (d) sequencing the third set of nucleic acid molecules or derivatives thereof. Optionally, the concentrating step includes contacting the second set of nucleic acid molecules or derivatives thereof with one or more nucleic acid capture probes. Optionally, the generating step includes contacting the first set of nucleic acid molecules or derivatives thereof with a methylated nucleic acid capture reagent. Optionally, the methylated nucleic acid capture reagent is formed by incubating methylated molecules with a solid substrate. Optionally, the solid substrate is beads. Optionally, the solid substrate is a magnetic solid substrate. Optionally, the solid substrate contains protein A. In some cases, the solid substrate contains streptavidin. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule contains biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the method or system further includes a step of amplifying a second set of molecules. In some cases, the amplification step is carried out while a subset of nucleic acids from the first set is bound to a methylated nucleic acid capture reagent. In some cases, the amplification is via PCR amplification. In some cases, the PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, the PCR amplification is 14 cycles of amplification. In some cases, the PCR amplification is 13 cycles of amplification. In some cases, the amplicons produced from the amplification may undergo cleanup. In some cases, the sequencing is sequencing by a synthetic reaction.In some cases, sequencing does not include bisulfite sequencing. In some cases, the method or system includes a step of performing one or more library preparation reactions on a third set of nucleic acid molecules before sequencing. In some cases, the method or system further includes a step of incubating the third set of nucleic acid molecules with a plurality of magnetic beads that interact with the nucleic acid after performing one or more library preparation reactions and before sequencing. In some cases, the method or system further includes a step of subjecting the nucleic acid sample to magnetic capture to remove the plurality of magnetic beads that interact with the nucleic acid after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid. In some cases, the method or system further includes a step of performing additional magnetic capture. In some cases, the method or system further includes a step of adding a certain amount of filler DNA to the first nucleic acid molecule before (b).
[0082]
[0109] In any of the methods or systems for processing nucleic acids described herein, quality control analysis may be performed on the captured methylated nucleic acids. For example, the quality control analysis may include measuring methylation bond specificity. In some cases, methylation bond specificity is measured by calculating the recovery rate of exogenous methylated fragments. In some cases, methylation bond specificity is measured by calculating the events of the detected methylated fragments. In some cases, methods and systems for processing nucleic acid samples from subjects described herein result in enrichment of methylated single-stranded DNA with methylation specificity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, or at least about 99.9%. In some cases, methods and systems for processing nucleic acid samples from subjects described herein result in an enrichment of methylated nucleic acids of at least about 1x, at least about 2x, at least about 3x, at least about 4x, at least about 5x, at least about 6x, at least about 7x, at least about 8x, at least about 9x, at least about 10x, at least about 15x, at least about 20x, at least about 25x, at least about 30x, at least about 35x, at least about 40x, at least about 45x, at least about 50x, at least about 55x, at least about 60x, at least about 65x, at least about 70x, at least about 75x, at least about 80x, at least about 85x, at least about 90x, at least about 95x, at least about 100x, at least about 150x, or at least about 200x.
[0083] Methylation profile
[0110] This disclosure provides methods and systems for generating methylation profiles of subjects having or suspected to have a disease / condition, which can be used to determine whether the subject has or is at risk of having such a disease / condition. In some cases, the methylation profile can be used to determine whether the subject has or is at risk of recurring a disease / condition (e.g., cancer). In some cases, the methylation profile can be used to determine whether the subject has not or is unlikely to recur a disease / condition (e.g., cancer). In some cases, the methylation profile may include analysis (e.g., sequencing) of multiple nucleic acids (e.g., multiple nucleic acid molecules in a depleted sequencing library, as described herein). In some cases, the methylation profile may include detection of methylated nucleotides and / or quantification of methylated nucleotide counts. In some cases, the methylation profile may include quantification of circulating tumor DNA (ctDNA). In some cases, ctDNA can be quantified over time (e.g., ctDNA dynamics). In some cases, methylation profiles may include determining methylation signals in a population of nucleic acids from a depleted sequencing library, for example, as described herein. In some cases, methylation profiles are compared to genome-wide background profiles. In some cases, methylation profiles are compared to novel background profiles created using highly methylated cfDNA.
[0084]
[0111] In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls) by an area under the receiver operating characteristic curve (AUROC) of at least approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or at least approximately 99.9%. In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls) with AUROCs of up to approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or up to approximately 99.9%. In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls) with AUROCs of approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%.
[0085] Genomic mutation profiles
[0112] This disclosure provides methods, systems, and kits for generating mutation profiles of subjects who have or are suspected of having a disease / condition, the mutation profiles of which may be used to determine whether the subject has or is at risk of having a disease / condition. Samples disclosed herein can be subjected to library preparation and next-generation deep sequencing to depths of, for example, 1 million (M) to 60M single reads, 10M to 60M single reads, 10M to 100M single reads, 40M to 60M single reads, 40M to 100M single reads, 60M to 100M single reads, 60M to 200M single reads, 1M to 10M single reads, 1M to 40M single reads, 1M to 100M single reads, 1M to 200M single reads, at least 1M single reads, at least 10M single reads, at least 40M single reads, at least 60M single reads, at least 100M single reads, or at least 200M single reads. In some cases, sequencing can be performed at low sequencing depths (e.g., single reads of 10M, 20M, 30M, 40M, 1M to 10M, 10M to 20M, 20M to 30M, 30M to 40M, up to 10M, up to 20M, up to 30M, or up to 40M).In some cases, the samples disclosed herein include 0.1×~100×, 0.1×~60×, 0.1×~40×, 0.1×~30×, 0.1×~20×, 0.1×~10×, 0.1×~5.0×, 0.5×~100×, 0.5×~60×, 0.5×~40×, 0.5×~30×, 0.5×~20×, 0.5×~10×, 0.5×~5.0×, 1.0×~100×, 1.0×~60×, 1.0×~40×, 1.0×~30×, 1.0×~20×, 1.0×~10×, 1.0×~5.0×, at least 0.1×, at least 0.5×, at least 1.0×, at least 2.0×, and at least All can be subjected to sequencing at depths of 3.0×, at least 4.0×, at least 5.0×, at least 10.0×, at least 20.0×, at least 30.0×, at least 40.0×, at least 50.0×, at least 60.0×, at least 100×, at least 200×, up to 0.1×, up to 0.5×, up to 1.0×, up to 2.0×, up to 3.0×, up to 4.0×, up to 5.0×, up to 10.0×, up to 20.0×, up to 30.0×, up to 40.0×, up to 50.0×, up to 60.0×, up to 100×, or up to 200×. Multiple sequencing reads are generated and analyzed. In some embodiments, deep sequencing may be configured to maximize the identification of disease / condition-related genomic variations.
[0086]
[0113] In some embodiments, the relative measure of ctDNA abundance is calculated from the mean variant allele proportion (MAF). In some embodiments, the mean MAF of mutations identified in a subject and included in the subject's mutation profile is in the range of at least about 0.01% to at least about 10%. In some cases, the MAF of the ctDNA fraction of a sample may be at least about 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage in between.
[0087]
[0114] In some embodiments, the generated mutation profile of a subject can be generated from sequencing results. In some embodiments, the mutation profile includes genetic polymorphisms such as missense variants, nonsense variants, deletion variants, insertion variants, duplication variants, inverted variants, frameshift variants, or repeat-extension variants. In some embodiments, the mutation profile may include mutation variants derived from a fraction of cell-free nucleic acid molecules within a specific size range. This disclosure provides methods, systems, and kits for generating mutation profiles of subjects who have or are suspected of having such a disease / condition, and methylation profiles may be used to determine whether the subject has or is at risk of having the disease / condition. The step of generating a genomic mutation profile may include the step of subjecting multiple nucleic acid molecules to library preparation and next-generation deep sequencing (e.g., MeDIP-seq). Multiple sequencing reads can be generated and analyzed, and in some cases, the deep sequencing may be configured to maximize the identification of genomic mutations associated with the disease / condition. For example, a panel of known cancer driver genes may be included in the selector for sequencing result analysis. In some embodiments, including genes that do not have a recorded driver effect in a particular type of cancer in the analysis of sequencing data may increase the sensitivity of ctDNA detection.
[0088]
[0115] In some embodiments, the relative measure of ctDNA abundance is calculated from the mean variant allele proportion (MAF). In some embodiments, the mean MAF of mutations identified in a subject and included in the subject's mutation profile is in the range of at least about 0.01% to at least about 10%. The ctDNA fractions of the samples disclosed herein are at least about 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage in between.
[0089]
[0116] In some embodiments, the generated mutation profile of the subject does not include mutation variants derived from cell-free nucleic acid molecules from the biological sample. In some embodiments, the mutation profile includes genetic polymorphisms such as missense variants, nonsense variants, deletion variants, insertion variants, duplication variants, inverted variants, frameshift variants, or repeating extension variants. In some embodiments, the mutation profile may include mutation variants derived from fractions of cell-free nucleic acid molecules within a specific size range.
[0090] Fragment length profile
[0117] In some embodiments, the length of the ctDNA fragment is shorter than that of a cell-free nucleic acid molecule derived from a healthy subject. In some embodiments, the length of the ctDNA containing at least one mutation is shorter than that of a cell-free nucleic acid molecule containing the corresponding reference allele.
[0091]
[0118] In some embodiments, sequencing does not utilize bisulfite sequences because they cause degradation of ctDNA fragments and interfere with the preservation of the ctDNA length distribution. In some embodiments, the fragment lengths of the multiple nucleic acids of this disclosure (e.g., those comprising a mixture of cfDNA molecules derived from tumor or cancerous tissue and healthy tissue, those comprising cfDNA molecules from healthy tissue only, and / or those comprising ctDNA only) may be 1 to about 800 base pairs (bp), about 50 bp to about 800 bp, about 100 bp to about 200 bp, about 120 bp to about 150 bp, about 60 to about 500 bp, about 80 to about 300 bp, 90 to about 250 bp, 80 to about 170 bp, or about 100 to about 150 bp. In some embodiments, the fragment lengths of the plurality of nucleic acids of the Disclosure (e.g., comprising a mixture of cfDNA molecules derived from tumor or cancerous tissue and healthy tissue, comprising cfDNA molecules from healthy tissue only, and / or comprising ctDNA only) may be at least 800 base pairs (bp), at least 700 base pairs, at least 600 base pairs, at least 500 base pairs, at least 400 base pairs, at least 300 base pairs, at least 200 base pairs, at least 150 base pairs, at least 100 base pairs, or at least 50 base pairs. In some embodiments, the fragment lengths of the multiple nucleic acids of this disclosure (e.g., those comprising a mixture of cfDNA molecules derived from tumor or cancerous tissue and healthy tissue, those comprising cfDNA molecules from healthy tissue only, and / or those comprising ctDNA only) may be up to 800 base pairs (bp), up to 700 base pairs, up to 600 base pairs, up to 500 base pairs, up to 400 base pairs, up to 300 base pairs, up to 200 base pairs, up to 150 base pairs, up to 100 base pairs, or up to 50 base pairs. In some embodiments, this disclosure provides enrichment of cell-free nucleic acid samples based on the selection of cell-free molecules of a particular size. In some embodiments, multimodal analysis includes utilizing the mutation profiles and fragment length profiles described herein by selectively including multiple nucleic acid molecules in the mutation profile based on their fragment lengths.In some embodiments, multimodal analysis includes utilizing the methylation profiles and fragment length profiles described herein by selectively including multiple nucleic acid molecules in the methylation profile based on their fragment lengths. In some embodiments, multimodal analysis includes utilizing the mutation profiles, methylation profiles, and fragment length profiles together by selectively including multiple nucleic acid molecules in the mutation profile based on their fragment lengths, and by selectively including multiple nucleic acid molecules in the methylation profile based on their fragment lengths, respectively.
[0092] Tumor detection and prognosis
[0119] This disclosure provides a method and system for determining whether a subject has or is at risk of having a disease, the method and system comprising the steps of: sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing the at least one profile to determine whether the subject has or is at risk of having the disease with at least 80% sensitivity or at least about 90% specificity, wherein the cell-free nucleic acid sample contains less than 30 ng / ml of the plurality of nucleic acid molecules. In some embodiments, the sensitivity is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. In some embodiments, the specificity is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers.
[0093]
[0120] In some embodiments, the method and system may include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least two profiles from among (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile. The method provides a sensitivity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. In some embodiments, the sensitivity when using two profiles increases by a percentage between at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any number, compared to the sensitivity when using one profile. In some embodiments, the sensitivity when using three profiles increases by a percentage between at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any number, compared to the sensitivity when using two profiles.
[0094]
[0121] Furthermore, this method can provide specificity of at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, or any percentage between these numbers. In some embodiments, the specificity when using two profiles increases by at least approximately 0.5%, at least approximately 1%, at least approximately 2%, at least approximately 3%, at least approximately 4%, at least approximately 5%, at least approximately 6%, at least approximately 7%, at least approximately 8%, at least approximately 9%, at least approximately 10%, or any percentage between these numbers, compared to the specificity when using one profile. In some embodiments, the specificity when using three profiles increases by at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or by a percentage between any number, compared to the specificity when using two profiles.
[0095]
[0122] This disclosure provides a method and system for processing a cell-free nucleic acid sample of a subject to determine whether the subject has or is at risk of having a disease, the method and system comprising the steps of: providing the cell-free nucleic acid sample comprising a plurality of nucleic acid molecules; subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads; processing the plurality of sequencing reads by computer to identify (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile for the plurality of nucleic acid molecules; and determining whether the subject has or is at risk of having the disease using at least the methylation profile, the mutation profile, and the fragment length profile. In some embodiments, the method provides a sensitivity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. This method provides specificity of at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, or any percentage between these numbers.
[0096]
[0123] This disclosure provides a method and system for determining the tissue origin of a tumor, comprising identifying nucleotide sequences specific to a particular cancer (e.g., breast cancer, colon cancer, prostate cancer, HSNCC, or lung cancer) from which a fraction of cell-free nucleic acid molecules originates. In some embodiments, the fraction of cell-free nucleic acid molecules originates from ctDNA. In some embodiments, the method provides a sensitivity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. This method provides specificity of at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, or any percentage between these numbers.
[0097]
[0124] This disclosure provides a method and system for determining whether a subject has or is at risk of having multiple diseases (e.g., multiple cancers), the method and system comprising the steps of: sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing the at least one profile to determine whether the subject has or is at risk of having the disease with at least 80% sensitivity or at least about 90% specificity, wherein the cell-free nucleic acid sample contains less than 30 ng / ml of the plurality of nucleic acid molecules. In some embodiments, the sensitivity is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. In some embodiments, specificity is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. In some cases, the method for determining whether a subject has or is at risk of having multiple diseases (e.g., multiple cancers) may include determining two, three, four, five, or more diseases.In some cases, when determining whether a subject has or is at risk of having multiple diseases (e.g., multiple cancers), it may be possible to identify no disease, a single disease (e.g., cancer), or multiple diseases (e.g., multiple cancers).
[0098]
[0125] This disclosure provides a method and system for determining whether a subject has cancer (e.g., low-exudative carcinoma) or is at risk thereof, the method and system comprising: (a) providing a plurality of nucleic acid molecules generated from a cfDNA sample of the subject; (b) subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads; (c) processing the plurality of sequencing reads by computer to generate methylation profiles of the plurality of nucleic acid molecules; and (d) processing the methylation profiles by computer to determine whether The process includes determining that the subject has cancer in an area under the receiver operating characteristic curve (AUROC) of at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, at least approximately 99.5%, at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, or at least approximately 99.9%. In some cases, AUROC may be up to approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%. In some cases, AUROC may be approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%.
[0099]
[0126] This disclosure provides a method and system for determining whether a subject has cancer (for example, endometrial cancer, esophageal cancer, hepatobiliary cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer), the method and system comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleotide (cfDNA) sample of the subject; (b) subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads in the absence of bisulfite conversion; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine whether the subject has cancer.
[0100]
[0127] This disclosure provides a method and system for determining whether a subject has experienced cancer recurrence, the method and system comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleotide (cfDNA) sample of the subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads in the absence of bisulfite conversion; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine whether the subject has cancer.
[0101]
[0128] This disclosure provides a method and system for determining whether a subject has a disease or is at risk of developing a disease, the method and system comprising the steps of: sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate a methylation profile; comparing the sequence of the captured cell-free methylated DNA with control cell-free methylated DNA sequences from healthy individuals and cancerous individuals; and detecting the disease of the subject by determining the statistically significant similarity between one or more sequences of the captured cell-free methylated DNA and cell-free methylated DNA sequences from cancerous individuals.
[0102]
[0129] In some embodiments, control cell-free methylated DNA sequences from healthy and cancerous individuals are included in a database of differential methylation regions (DMRs) between healthy and cancerous individuals.
[0103]
[0130] This disclosure describes methods and systems for providing prognosis to subjects after treatment for a disease / condition. For example, treatment includes surgical removal of a tumor, chemotherapy, radiotherapy, or immunotherapy (e.g., TCR, CAR, etc.) designed for a specific type of cancer. In some embodiments, the method or system includes the steps of sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile, and monitoring or detecting minimal residual disease (MRD) based at least on the at least one profile.
[0104]
[0131] After the patient has been accurately diagnosed and has received treatment for cancer, such as surgical resection, chemotherapy, or radiation therapy, it may be important to monitor the effectiveness of the treatment and predict the patient's survival rate. Furthermore, it may be important to detect minimal residual disease (MRP) of cancer cells.
[0105]
[0132] In some embodiments, the method further includes the step of adding a second amount of control DNA to the sample to confirm the immunoprecipitation reaction.
[0133] As used herein, “control” may include both positive and negative controls, or at least a positive control.
[0106]
[0134] In some embodiments, the method further includes the step of adding a second amount of control DNA to the sample to confirm the capture of cell-free methylated DNA.
[0135] In some embodiments, the step of identifying the presence of cancer cell-derived DNA further includes the step of identifying the originating cancer cell tissue.
[0107]
[0136] In some cases, tumor tissue sampling can be difficult or carry significant risks, in which case it may be desirable to diagnose and / or subtype cancer without requiring tumor tissue sampling. For example, lung tumor tissue sampling may require invasive procedures such as mediastinoscopy, thoracotomy, or percutaneous needle biopsy, which may necessitate hospitalization, pleural tube insertion, mechanical ventilation, antibiotics, or other medical interventions. Some individuals may not be able to undergo the invasive procedures required for tumor tissue sampling due to medical comorbidities or at their own request. In some cases, the actual procedure for obtaining tumor tissue may depend on the suspected cancer subtype. In other cases, cancer subtypes may develop over time within the same individual, and sequential evaluation by invasive tumor tissue sampling procedures is often impractical and poorly tolerated by patients. Therefore, non-invasive cancer subtype classification by blood tests has many beneficial applications in the practice of clinical oncology.
[0108]
[0137] Therefore, in some embodiments, the step of identifying the originating cancer cell tissue further includes the step of identifying the cancer subtype. In some cases, the cancer subtype distinguishes cancers based on stage (e.g., early lung cancer treated with surgery and late lung cancer treated with chemotherapy), histological examination (e.g., small cell carcinoma, adenocarcinoma and squamous cell carcinoma in lung cancer), gene expression pattern or transcription factor activity (e.g., ER status in breast cancer), copy number abnormalities (e.g., HER2 status in breast cancer), specific rearrangements (e.g., FLT3 in AML), specific gene point mutation status (e.g., IDH gene point mutation), and DNA methylation pattern (e.g., MGMT gene promoter methylation in brain cancer).
[0109]
[0138] In some embodiments, the comparison may be based on fitting using a statistical classifier. In some cases, a statistical classifier using DNA methylation data can be used to assign a sample to multiple cancers. In some cases, a statistical classifier using DNA methylation data can be used to assign a sample to multiple cancers for the detection of multiple cancers. In some cases, multiple cancer detection may be multicancerous early detection (MCED). In some cases, a statistical classifier using DNA methylation data can be used to assign a sample to minimal residual disease (MRD). In some cases, a statistical classifier using DNA methylation data can be used to assign a sample to a specific disease state, such as a cancer type or subtype (e.g., early cancer, low-exudative tumor). In some cases, if the classifier can distinguish multiple cancer types (or subtypes) from one another, the classifier may have differential methylation regions from pairwise comparisons of each cancer type (or subtype) of interest. In some cases, the statistical classifier may have one or more DNA methylation variables in the statistical model, and / or the output of the statistical model may have one or more thresholds for distinguishing different disease states. In some cases, the statistical classifier may have features and / or thresholds that can be derived from prior knowledge of cancer types or subtypes, prior knowledge of features that may be the most useful sources of information, machine learning, or a combination of two of these approaches. In some embodiments, the classifier may be derived from machine learning. In some cases, the classifier may be an elastic network classifier, Lasso, support vector machine, random forest, or neural network.
[0110]
[0139] In some embodiments, the comparison can be performed genome-wide. In other embodiments, the comparison can be limited from genome-wide to specific regulatory regions, including but not limited to long scattered repeats (LINEs), short scattered repeats (SINEs), terminal repeats (LTRs), FANTOM5 enhancers, CpG islands, CpG Shores, CpG shelves, or any combination thereof.
[0111]
[0140] In some embodiments, the methods described herein are for use in the detection of cancer. In some embodiments, the methods described herein are for use in the detection of early-stage cancer (e.g., early-stage cancer). In some embodiments, the methods described herein are for use in the detection of early-stage (e.g., early-stage cancer) multiple cancers. In some embodiments, the methods described herein are for use in the detection of low-exudative tumors. In some embodiments, the methods described herein are for use in the detection of MRD. In some embodiments, the methods described herein are for use in the detection of low ctDNA.
[0112]
[0141] In some embodiments, the methods described herein are intended for use in monitoring cancer treatment. Data analysis system and method
[0142] Methods and systems disclosed herein may include algorithms or the use thereof. One or more algorithms may be used to classify one or more samples from one or more subjects. One or more algorithms may be used to quantify ctDNA. One or more algorithms may be used to predict a response to treatment. One or more algorithms may be used to predict a disease / condition (e.g., cancer). One or more algorithms may be applied to data from one or more samples. The data may include biomarker expression data. In some embodiments, the method or system includes the steps of sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subjects to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile, and monitoring or detecting minimal residual disease (MRD) based on at least one profile. Methods disclosed herein may include assigning classifications to one or more samples from one or more subjects. Assigning classifications to samples may include applying algorithms to the methylation profile, mutation profile, and fragment length profile. In some cases, at least one profile is fed into a data analysis system that includes a trained algorithm for classifying samples as having originated from subjects with disease or minor injury.
[0113]
[0143] The data analysis system may be a trained algorithm. The algorithm may include a linear classifier. In some examples, the linear classifier may include one or more of the following: linear discriminant analysis, Fisher's linear discriminant analysis, a simple Bayesian classifier, logistic regression, a perceptron, a support vector machine, or a combination thereof. The linear classifier may be a support vector machine (SVM) algorithm. The algorithm may also include a binary classifier. The binary classifier may include one or more decision trees, random forests, Bayesian networks, support vector machines, neural networks, or logistic regression algorithms.
[0114]
[0144] The algorithm may include one or more linear discriminant analysis (LDA), basic perceptron, elastic network, logistic regression, (kernel) support vector machine (SVM), diagonal linear discriminant analysis (DLDA), Golub classifier, Parzen estimation, (kernel) Fisher discriminant classifier, k nearest neighbor, iterative RELIEF, classification tree, maximum likelihood classifier, random forest, nearest neighbor centroid, predictive analysis by microarray (PAM), k-median clustering, fuzzy C-means clustering, Gaussian mixture models, stepwise responses (GR), gradient boosting (GBM), elastic network logistic regression, logistic regression, or a combination thereof. The algorithm may include the diagonal linear discriminant analysis (DLDA) algorithm. The algorithm may include the nearest neighbor centroid algorithm. The algorithm may include the random forest algorithm. In some embodiments, logistic regression, random forests, and gradient boosting (GBM) perform better than linear discriminant analysis (LDA), neural networks, and support vector machines (SVM) for distinguishing between pre-eclampsia and non-eclampsia.
[0115]
[0145] This disclosure provides a method and system for determining whether a subject has or is at risk of having a disease, the method and system comprising the steps of: sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing the at least one profile to determine whether the subject has or is at risk of having the disease with at least 80% sensitivity or at least about 90% specificity, wherein the cell-free nucleic acid sample contains less than 30 ng / ml of the plurality of nucleic acid molecules. In some embodiments, the sensitivity is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the specificity is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers.
[0116]
[0146] In some embodiments, the method and system may include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least two profiles from among (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile. The method provides sensitivity of at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the sensitivity when using two profiles increases by at least approximately 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any of these percentages, compared to the sensitivity when using one profile. In some embodiments, the sensitivity when using three profiles increases by at least approximately 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any of these percentages, compared to the sensitivity when using two profiles.
[0117]
[0147] Furthermore, this method can provide specificity of at least approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the specificity when using two profiles increases by at least approximately 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers, compared to the specificity when using one profile. In some embodiments, the specificity when using three profiles increases by at least approximately 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers, compared to the specificity when using two profiles.
[0118]
[0148] This disclosure provides a method and system for processing a cell-free nucleic acid sample of a subject to determine whether the subject has or is at risk of having a disease, the method and system comprising the steps of: providing the cell-free nucleic acid sample comprising a plurality of nucleic acid molecules; subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads; processing the plurality of sequencing reads by computer to identify (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile for the plurality of nucleic acid molecules; and determining whether the subject has or is at risk of having the disease using at least the methylation profile, the mutation profile, and the fragment length profile. In some embodiments, the method provides sensitivity of at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. The method can provide specificity of at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers.
[0119]
[0149] This disclosure describes methods and systems for providing prognosis to subjects after treatment for a disease / condition. For example, treatment includes surgical removal of a tumor, chemotherapy, radiotherapy, or immunotherapy (e.g., TCR, CAR, etc.) designed for a specific type of cancer. In some embodiments, the method or system includes the steps of sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one profile of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile, and monitoring or detecting minimal residual disease (MRD) based on at least one profile.
[0120] Computer system
[0150] This disclosure provides a computer system programmed to carry out the method of this disclosure. Figure 2 shows a computer system 201 programmed, or otherwise configured, to generate a sequencing library containing nucleic acid molecules (e.g., ctDNA) with depleted hypermethylation regions. The computer system 201 can be adapted to various aspects of this disclosure. The computer system 201 may be a computer system located remotely from a user's electronic device. The electronic device may be a mobile electronic device.
[0121]
[0151] The computer system 201 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes memory or memory locations 210 (e.g., random-access memory, read-only memory, flash memory), electronic storage devices 215 (e.g., hard disks), a communication interface 220 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 225 such as caches, other memory, data storage, and / or electronic display adapters. The memory 210, storage devices 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a communication bus (solid line), such as a motherboard. The storage device 215 may be a data storage device (or data repository) for storing data. The computer system 201 can be operably coupled to a computer network (“network”) 230 using the communication interface 220. Network 230 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. Network 230 may optionally be a telecommunications and / or data network. Network 230 may include one or more computer servers that can enable distributed computing, such as cloud computing. Network 230 may optionally implement a peer-to-peer network using computer system 201, which can enable devices coupled to computer system 201 to act as clients or servers.
[0122]
[0152] The CPU 205 can execute a sequence of machine-readable instructions that may be embodied in a program or software. Instructions can be stored in a memory location, such as memory 210. Instructions can be directed to the CPU 205, which can then be programmed or otherwise configured to perform the methods of this disclosure. Examples of operations performed by the CPU 205 include fetching, decoding, executing, and writing back.
[0123]
[0153] The CPU 205 may be part of a circuit, such as an integrated circuit. One or more other components of system 201 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0124]
[0154] The storage device 215 can store files such as drivers, libraries, and saved programs. The storage device 215 can also store user data, such as user preferences and user programs. The computer system 201 may, in some cases, include one or more additional data storage devices located outside the computer system 201, such as those located on a remote server that communicates with the computer system 201 via an intranet or the internet.
[0125]
[0155] Computer system 201 can communicate with one or more remote computer systems via network 230. For example, computer system 201 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 201 via network 230.
[0126]
[0156] The methods described herein can be implemented by machine-executable code (e.g., a computer processor) stored in an electronic storage location of the computer system 201, such as memory 210 or electronic storage device 215. The machine-executable code or machine-readable code can be provided in the form of software. In use, the code can be executed by the processor 205. In some cases, the code can be retrieved from the storage device 215 and stored in memory 210 for easy access by the processor 205. In some situations, the electronic storage device 215 may be omitted, and the machine-executable instructions are stored in memory 210.
[0127]
[0157] The code may be pre-compiled and configured for use on a machine with a processor adapted to run the code, or it may be compiled during runtime. The code may be supplied in a programming language that can be chosen to allow the code to run in a pre-compiled form or in an as-compiled form.
[0128]
[0158] In some embodiments, the computer system may include computer processing using supervised or unsupervised machine learning methods. In some cases, the supervised machine learning method may be regression, support vector machines, tree-based methods, neural networks, or nearest neighbor methods. In some cases, the unsupervised machine learning method may be clustering, neural networks, principal component analysis, or matrix factorization.
[0129]
[0159] Embodiments of systems and methods provided herein, such as computer system 201, can be embodied in programming. Various embodiments of the art can typically be considered “products” or “manufactured products” in the form of machine (or processor) executable code and / or associated data contained on or embodied within some kind of machine-readable medium. Machine-executable code can be stored in electronic storage devices such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disks. “Storage” type media can include any or all of tangible memory such as computers, processors, or their associated modules, such as various semiconductor memory, tape drives, disk drives, etc., which can provide non-temporary storage for software programming at any time. All or part of the software may be communicated over the Internet or various other telecommunication networks. Such communication can enable, for example, the loading of software from one computer or processor to another computer or processor, for example, from a management server or host computer to an application server computer platform. Therefore, other types of media that can hold software elements include optical waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices through wired and optical landline network connections and various air links. Physical elements that carry such waves, such as wired or wireless links and optical links, can also be considered media containing software. Unless limited to non-temporary tangible “storage” media as used herein, the terms such as computer or machine “readable media” refer to any medium involved in providing instructions to a processor for execution.
[0130]
[0160] Therefore, machine-readable media such as computer executable code can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any storage device in any computer, which can be used to implement, for example, a database as shown in the drawing. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wires, and optical fibers, including wires including buses in computer systems. Carrier media can take the form of electrical or electromagnetic signals, or sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, perforated tapes, any other physical storage media having a pattern of holes, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers that carry data or instructions, cables or links that carry such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0131]
[0161] The computer system 201 includes, or can communicate with, an electronic display 1135 equipped with a user interface (UI) 240. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0132]
[0162] The methods and systems of this disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software during execution by the central processing unit 205.
[0133]
[0163] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided only as examples. The present invention is not intended to be limited by any specific examples provided herein. The present invention has been described with reference to the preceding specification, but the descriptions and examples of embodiments herein are not intended to be constrained. Those skilled in the art will be able to conceive of numerous variations, modifications, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to any specific descriptions, configurations, or relative proportions described herein, depending on various conditions and variables. It should be understood that various alternative forms to the embodiments of the present invention described herein may be used when carrying out the present invention. Accordingly, the present invention is intended to encompass any such alternative forms, modifications, variations, or equivalents. The following claims define the scope of the present invention, and the methods and structures within these claims, as well as their equivalents, are intended to be encompassed thereby.
[0134] kit
[0164] This disclosure provides a kit for identifying or monitoring a disease or disorder of interest (e.g., cancer). The kit may include probes for identifying quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences in each of a panel of cancer-related genomic loci in a sample of interest. These quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences in each of the panel of cancer-related genomic loci in a sample may indicate a disease or disorder of interest (e.g., cancer). The probes may be selective for sequences in the panel of cancer-related genomic loci in a sample. The kit may include instructions for processing a sample using the probes and generating a dataset showing quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences in each of the panel of cancer-related genomic loci in a sample of interest.
[0135]
[0165] The probes in the kit may be selective for sequences in a panel of cancer-related genomic loci in the sample. The probes in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to the panel of cancer-related genomic loci. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with one or more nucleic acid sequences from a panel of cancer-related genomic loci or genomic regions. The panel of cancer-related genomic loci or microbiome-related genomic loci or genomic regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more separate panels of cancer-related genomic loci or genomic regions.
[0136]
[0166] The instructions included in the kit may include instructions for assaying a cell-free biological sample using probes that are selective for sequences in a panel of cancer-related genomic loci in the sample. These probes may be nucleic acid molecules (e.g., RNA or DNA) that are sequence-complementary to nucleic acid sequences (e.g., RNA or DNA) from one or more of the panels of cancer-related genomic loci. These nucleic acid molecules may be primers or enriched sequences. The instructions for assaying a cell-free biological sample may include an introduction to performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the sample and generate a dataset showing quantitative measures (e.g., presence, absence, or relative amount) of sequences in each of the panels of cancer-related genomic loci in the sample. Quantitative measures (e.g., presence, absence, or relative amount) of sequences in each of the panels of cancer-related genomic loci in the sample may indicate a disease or disorder (e.g., cancer).
[0137]
[0167] The instructions included in the kit may include instructions for measuring and interpreting assay readouts, which may be quantified in one or more panels of cancer-related genomic loci to generate a dataset showing a quantitative measure (e.g., presence, absence, or relative amount) of the sequence in each panel of cancer-related genomic loci in the sample. For example, quantification by array hybridization or polymerase chain reaction (PCR) corresponding to a panel of cancer-related genomic loci may generate a dataset showing a quantitative measure (e.g., presence, absence, or relative amount) of the sequence in each panel of cancer-related genomic loci in the sample. Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof. [Examples]
[0138] Example 1: Processing of plasma-derived cell-free DNA using a total methylome enrichment platform
[0168] Cell-free methylated DNA immunoprecipitation and high-through-out sequencing (cfMeDIP-seq) was developed as a non-degradable liquid biopsy approach to circumvent the limitations of bisulfite sequencing. The bisulfite-free approach may require consistently high methylation-binding specificity to detect potentially rare methylation events in circulating tumor DNA (ctDNA). Furthermore, as shown in Figure 23, as genomic DNA contamination increased, the effective cfDNA input to the cfMeDIP-seq workflow decreased, thereby reducing the probability of detecting rare ctDNA and highlighting the need for workflow improvements. gDNA contamination was measured by electrophoretic fragment size profiling (e.g., TapeStation). No correlation was observed between methylation specificity and genomic DNA contamination rates. Figure 24 also illustrates the importance of high methylation specificity for methylation-based liquid biopsy assays. This figure compares two clinical samples with high (y-axis) and low (x-axis) methylation specificity. A CpG count of 0 indicates the count of sequencing fragments aligned with regions in the human genome where the known CpG count is 0. The slope in the figure shows the number of non-CpG regions in the genome with observed counts for each sample. Reads without CpG and deviations from 100% methylation specificity percentage indicate nonspecific binding by the anti-5mC antibody during immunoprecipitation. Perfect methylation specificity (i.e., 100%) in immunoprecipitation results in no detection of DNA fragment counts in non-CpG regions. In real samples, generally, lower methylation specificity leads to more counts being observed in CpG-free regions. Therefore, in this example, cfMEDIP-seq was further improved with the aim of developing a robust genome-wide methylome enrichment platform for clinical use.
[0139]
[0169] As shown in Figure 25, four different cfMeDIP-seq workflows were performed using cell-free DNA (cfDNA) samples, and the methylation specificity of each workflow was determined. Samples included those derived from the target plasma, as well as samples derived from sheared genomic DNA to mimic cfDNA. For workflow 1, cfDNA samples (n=110) mixed with spike-in DNA were subjected to library preparation. Next, filler DNA was added to the prepared library to create a sample mixture. The sample mixture was thermally denatured and rapidly cooled. Then, 5-mc binder and magnetic beads were sequentially added to the mixture and incubated to bind single-stranded DNA (ssDNA) to the magnetic bead-5-mc binder complex. Using a magnet, the magnetic bead-5-mc binder complex was isolated, and the captured ssDNA was eluted from the complex. The captured ssDNA was cleaned up and subjected to 14 cycles of PCR amplification. The PCR amplicon was cleaned up again before quality control analyses such as measurement of methylation specificity.
[0140]
[0170] Workflow 2 (n=174) was performed similarly to Workflow 1, but with the addition of a bead capture step before adding filler DNA. To perform bead capture, the eluate from the cleanup after library preparation was subjected to further capture using a magnetic rack, and all beads carried over from the previous step (e.g., solid-phase reversible immobilization (SPRI) beads) were discarded to minimize any nonspecific binding to the beads. Workflow 3 (n=63) was performed similarly to Workflow 2, except that the number of cycles used in the PCR amplification step was reduced to 13 cycles.
[0141]
[0171] Workflow 4 (n=189) followed a similar protocol to Workflow 3, except that a pre-binding step was performed to prepare a magnetic bead-5-mC binder complex by separately mixing the magnetic beads and 5-mC binder. After thermal denaturation and rapid cooling of the mixture sample, the magnetic bead-5-mC binder complex was added. Another difference between Workflow 4 and Workflow 3 was that instead of eluting the captured ssDNA from the magnetic bead-5-mC binder complex and washing the captured ssDNA for PCR amplification, the PCR amplification step was performed on the beads.
[0142]
[0172] As shown in Figure 26, mean methylation specificity was measured for different workflows by calculating the percentage of synthetic oligonucleotides with known methylation states spiked to sample DNA during the library preparation step, which was recovered by immunoprecipitation (e.g., recovery of exogenous methylated fragments). Compared to the other workflows, workflow 4 showed higher methylation specificity and reduced sample-to-sample variability.
[0143]
[0173] The methylation specificity of the concentration platform was also analyzed for samples with different tube and sample storage periods. Table 1 shows the effect of tube and sample storage periods on methylation specificity.
[0144] [Table 1]
[0145]
[0174] The results showed that methylation specificity was consistent regardless of the type of tube (Δ=0.31%) or the storage period of the sample (Δ<0.26%). Example 2: Genome-wide methylome enrichment platform for early detection of multiple cancers (MCED)
[0175] Genome-wide mapping of DNA methylation in circulating cell-free DNA (cfDNA) can overcome a significant sensitivity problem in detecting circulating tumor DNA (ctDNA) in subjects with early-stage cancer or low-exudative tumors.
[0146]
[0176] To investigate a genome-wide methylome enrichment platform for the early detection of multiple cancers, a retrospective case-control study was conducted using plasma samples from commercial biobanks and the University Health Network biobank. Samples were taken from individuals diagnosed with cancer but who had not initiated treatment. Non-cancer controls were obtained from age- and sex-matched individuals with no known history of cancer diagnosis. Controls had a cancer-free follow-up period of at least 12 months after sampling. Individuals who were 75 years of age or older at the time of sampling, or who were known to have multiple comorbidities, were excluded from the control group.
[0147]
[0177] 5–10 ng of cfDNA was extracted from plasma samples and analyzed on a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) techniques, as described in Example 1, Figure 25 (Workflow 1). To distinguish cases from controls, samples were split into separate sets and a machine learning classifier (e.g., a linear-based approach) consisting of differentially methylated regions across the entire methylome was trained and tested. Initial training was performed with 1,536 samples across eight cancer types, followed by cross-validation with a random split (80:20) of 100 iterations to obtain the area under the receiver operating characteristic curve (AUC) and its 95% confidence interval (CI) using the median of the predicted probability. As shown in Table 2, all cancer cases were distinguished from controls with an AUC of 0.94 (95% CI: 0.93, 0.96), and the AUC for individual cancer types (e.g., bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, ovarian cancer, prostate cancer, kidney cancer) ranged from 0.91 to 0.97. The AUC for all cancers across all stages was 0.94 (95% CI: 0.92, 0.95) for stage I / II cancers and 0.95 (95% CI: 0.94, 0.96) for stage III / IV cancers. The AUC was 0.92 (95% CI: 0.91, 0.94) for a subset of low-exudative cancers (e.g., bladder cancer, breast cancer, prostate cancer, and kidney cancer), and showed similar performance for stages I / II (AUC 0.91; 95% CI: 0.89, 0.93) and stages III / IV (AUC 0.93; 95% CI: 0.91, 0.95). These results suggest that a genome-wide methylome enrichment platform can detect early-stage, multiple, and low-exudative cancer subsets.
[0148] [Table 2]
[0149]
[0178] Preliminary training was also performed on 1,906 samples across 12 types of cancer (e.g., bladder cancer, breast cancer, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, and kidney cancer). Table 3 shows the characteristics between cancer cases and control cases. A machine learning classifier was used to distinguish cancer cases from non-cancer controls. Cross-validation was performed with 20 iterations, five times the dataset, to obtain AUC and 95% CI using the median of the predicted probability.
[0150] [Table 3]
[0151]
[0179] As shown in Figure 3, the genome-wide methylation assay distinguished cancer cases from non-cancer controls with an overall AUC of 0.94 (95% CI: 0.93, 0.95). Stage-specific AUCs were 0.92 (Stage I), 0.95 (Stage II), 0.95 (Stage III), and 0.97 (Stage IV), indicating that AUC increased with cancer stage but remained high even in early-stage cancers. When each cancer type was evaluated separately (e.g., bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer), the AUC ranged from 0.89 to 0.99, as shown in Figures 4A–4L. Subsets of cancers with low exudative tumors (e.g., bladder cancer, breast cancer, endometrial cancer, and kidney cancer) were also evaluated separately. The AUC of this subset of low-exudative tumors was 0.91 (95% CI: 0.89, 0.93), and as shown in Table 4, only 8% were considered to be stage IV cancers, suggesting that MCED is possible using a genome-wide methylome enrichment platform.
[0152] [Table 4]
[0153] Example 3: Evaluation of a genome-wide methylome enrichment platform for ctDNA quantification in renal cell carcinoma (RCC).
[0180] Circulating tumor DNA (ctDNA) can be used to identify the presence of cancer and minimal residual disease. Quantifying ctDNA using plasma-based tests may be a useful cancer management tool for assessing prognosis; however, some methods also require tumor tissue for analysis or are limited to tumor types that tend to have relatively high amounts of associated ctDNA. This experiment demonstrated the potential to quantify ctDNA in plasma and predict prognosis in recurrent cancer (RCC) using a genome-wide methylome enrichment platform independent of tumor information.
[0154]
[0181] Pre-treatment samples stored in biobanks (University Health Network, Ontario Tumor Biobank) from newly diagnosed individuals with stage I–IV RCC were analyzed using 5–10 ng of cell-free DNA isolated from plasma on a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform, as described in Example 1, Figure 25 (Workflow 1). ctDNA was quantified from the mean-normalized count across the entire information region.
[0155]
[0182] The algorithm identified cancer-associated methylation using reduced-size fragments of 1–150 base pairs (bp) containing at least 10 CpGs. Counts were aggregated over non-overlapping 300 bp windows and normalized by sequencing depth. Methylation signals were contrasted across each 300 bp window using a single factor (e.g., cancer vs. non-cancer control). For each window, p-values were calculated between cancer and non-cancer control using the Wilcoxon rank-sum test. P-values were adjusted using the Benjamini-Hochberg method. Significant differential methylation regions were selected using an adjusted p-value threshold of ≤0.1. Differential methylation analysis identified 2027 hypermethylated regions where CpG islands were enriched. ctDNA quantification scores were generated based on mean-normalized counts across the 2027 regions and corrected for methylation specificity. Events were defined as cancer recurrence or progression. The ctDNA threshold was set so that 95% of non-event samples were below the threshold (e.g., 95% specificity). For samples with ctDNA levels above the threshold, the time to event occurrence was compared with samples below the threshold. Figure 5A shows that the threshold for renal cancer was set at 0.37.
[0156]
[0183] The cohort included 151 samples [64 stage I cases, 2 stage II cases, 23 stage III cases, 15 stage IV cases, and 47 cases with unknown or incomplete stage information]. The median follow-up period was 15.7 months, during which a total of 21 events occurred. Samples with ctDNA levels above the threshold were significantly more likely to recur or progress than those with levels below the threshold [hazard ratio 13.28 (95% CI 5.47, 32.26), log-rank P<0.001] (Figure 5B). Samples with ctDNA levels below the threshold were more likely to avoid cancer recurrence or relapse.
[0157]
[0184] This experiment demonstrated the potential of a blood-based, tumor-independent, genome-wide methylome enrichment platform for ctDNA quantification and prognosis prediction in renal cancer. This is a promising demonstration of prognostic performance in a cancer type typically difficult to detect due to low levels of ctDNA. Further evaluation in post-treatment and longitudinal samples will test this platform further.
[0158]
[0185] Another similar study included plasma samples stored in biobanks from individuals with newly diagnosed stage I–IV RCC (collected between 2015 and 2021; from the Princess Margaret Cancer Centre within the University Health Network and the Ontario Tumors Bank). All samples were obtained after cancer diagnosis but before surgery or other final treatment. Table 5 shows the clinicopathological information for the 148 samples used, and Figure 8 shows the age of the individuals at the time of sample collection.
[0159]
[0186]
[0160] [Table 5]
[0161]
[0187] All samples were analyzed using approximately 5–10 ng of cfDNA extracted from plasma on a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform. The genome-wide methylation assays used were based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) techniques, as described in Example 1, Figure 25 (Workflow 1).
[0162]
[0188] For analysis, an event was defined as cancer recurrence, progression, or death from renal cancer (whichever occurred earlier). The ctDNA dose threshold for baseline prognosis was set so that approximately 95% of the event-free samples were below the threshold (i.e., approximately 95% specificity). Event-free survival was estimated using the Kaplan-Meier method, and the difference was assessed by log-rank tests using both the whole population and individuals with stage I–III cancer.
[0163]
[0189] As shown in Figure 9, individuals with ctDNA quantification exceeding the cutoff value had significantly worse event-free survival in the entire cohort. Furthermore, as shown in Figure 10, individuals with ctDNA quantification exceeding the cutoff also had significantly worse event-free survival in the subpopulation with stage I-III disease. Individuals were stratified based on whether their ctDNA quantification was above or below the 95% specificity cutoff.
[0164]
[0190] Therefore, these data demonstrate the potential of using a blood-based, genome-wide methylome enrichment platform for ctDNA quantification to determine prognostic performance in RCCs. The observed performance also provides a promising demonstration of prognostic determination in cancer types that are typically difficult to detect due to low levels of ctDNA. Furthermore, the assay utilized here is tumor information-independent, meaning that patient-specific tumor tissue is not required to generate a tailored panel of ctDNA.
[0165] Example 4: Prognostic performance of a genome-wide methylome enrichment platform in head and neck cancer
[0191] The use of plasma-based tests to quantify circulating tumor DNA (ctDNA) is emerging as a promising new approach to cancer management. ctDNA quantification can be used to assess prognosis and detect minimal residual disease after initial treatment. This experiment demonstrated the potential of using a tumor-independent, genome-wide methylome enrichment platform for ctDNA quantification and prognosis prediction in head and neck cancer.
[0166]
[0192] Pre-treatment samples stored in a biobank (University Health Network) from individuals with newly diagnosed stage I–IV head and neck cancer were analyzed using a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform with 5–10 ng of cell-free DNA isolated from plasma, as described in Example 1, Figure 25 (Workflow 4). ctDNA was quantified from the mean-normalized count across the entire information region.
[0167]
[0193] The algorithm identified cancer-associated methylation using reduced-size fragments of 1–150 bp containing at least 10 CpGs. Counts were aggregated across non-overlapping 300 bp windows and normalized by sequencing depth. Methylation signals were contrasted across each 300 bp window using a single factor (e.g., cancer vs. non-cancer control). For each window, p-values were calculated between cancer and non-cancer control using the Wilcoxon rank-sum test. P-values were adjusted using the Benjamini-Hochberg method. Significant differential methylation regions were selected using an adjusted p-value threshold of ≤0.1. Differential methylation analysis identified 2027 hypermethylated regions where CpG islands were enriched. ctDNA quantification scores were generated based on mean-normalized counts across the 2027 regions and corrected for methylation specificity. Events were defined as cancer recurrence or progression. The ctDNA level threshold was set so that 95% of non-event samples were below the threshold (e.g., 95% specificity). For samples with ctDNA levels above the threshold, the time to recurrence or progression events was compared with samples below the threshold. Figure 6A shows that the threshold for head and neck cancer was set at 3.96.
[0168]
[0194] Of the 93 samples included (7 stage I, 17 stage II, 23 stage III, and 46 stage IV), the median follow-up period was 50.6 months, and 25 events occurred. The likelihood of recurrence or progression was significantly higher in samples with ctDNA above the threshold [hazard ratio (HR) 3.18 (95% CI 1.09, 9.28), log-rank P=0.026] (Figure 6B). In a multivariate analysis considering cancer stage and clinical characteristics (sex, age, smoking history, BMI), ctDNA levels above the threshold showed a similar association [HR 3.51 (95% CI 1.1, 11.19), P=0.034]. This experiment demonstrated the potential of using a genome-wide methylome enrichment platform, independent of blood-based tumor information, for ctDNA quantification and prognosis prediction in head and neck cancer, using plasma samples from treatment-naive individuals.
[0169]
[0195] Furthermore, the use of plasma-based tests to quantify ctDNA can also be used to assess prognosis and improve post-treatment monitoring. For head and neck cancer, this test provides information to help determine whether chemotherapy should be administered after surgery in some non-metastatic tumors (stages I-IVb), and offers an opportunity to improve post-treatment monitoring by detecting minimal residual disease before clinical detection.
[0170]
[0196] Another similar study included plasma samples from individuals with newly diagnosed stage I–IV head and neck cancer (collected between 2008 and 2019; Princess Margaret Cancer Centre, University Health Network). All samples were obtained after cancer diagnosis but before surgery or other definitive treatment. All samples were analyzed using 5–10 ng of cell-free DNA isolated from the plasma on a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform. This assay was based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) techniques, as described in Example 1, Figure 25 (Workflow 4). ctDNA was quantified across the entire information region using machine learning algorithms.
[0171]
[0197] For analysis, an event was defined as cancer recurrence, progression, or death from head and neck cancer (whichever occurred earlier). The maximum follow-up period was 5 years after diagnosis. The ctDNA dose threshold was set so that 95% of event-free samples were below the threshold (i.e., 95% specificity). Event-free survival was estimated using the Kaplan-Meier method and compared between samples with ctDNA above the threshold and samples with ctDNA below the threshold. The difference between the two groups was assessed by log-rank tests. Multivariate Cox regression analysis was used to adjust for known prognostic covariates.
[0172]
[0198] Table 6 shows the clinicopathological information of the subjects. Table 7 shows an overview of the first-line treatment for the cancers in the subjects. As shown in Tables 6 and 7, a total of 91 samples were included, with an average follow-up period of 50.6 months and 27 events occurring.
[0173]
[0199] Individuals with ctDNA levels above the threshold had significantly worse event-free survival, as shown in Figure 7. Lead times (time interval from blood collection to event) ranged from 1.25 to 16.37 months, with a median lead time of 4.87 months (Figure 7). In multivariate analysis based on cancer stage and clinical characteristics, ctDNA levels above the threshold showed a similar association (Table 8).
[0174] [Table 6]
[0175] [Table 7]
[0176] [Table 8]
[0177]
[0200] Similar to previous experiments, this study also demonstrated the potential of using a blood-based tumor information-independent, genome-wide methylome enrichment platform for ctDNA quantification and prognosis prediction in head and neck cancers. This study utilized a tumor information-independent method, meaning that patient-specific tumor tissue is not required to generate a customized panel for ctDNA detection. The tumor information-independent approach allows for more flexible clinical use by 1) not requiring access to original tumor tissue and 2) not being limited to differentially methylated genomic regions at the time of diagnosis, potentially enabling highly sensitive longitudinal minimal residual disease (MRD) detection.
[0178]
[0201] A tissue-independent, genome-wide methylome enrichment platform based on cfMEDIP-seq in head and neck cancer can be used to predict recurrence and detect early recurrence, with the aim of guiding adjuvant therapy after completion of curative treatment. To investigate the detection of MRD in head and neck cancer patients after curative treatment, longitudinal data collection and sampling (Princess Margaret Cancer Centre) were performed, and biobank-stored samples from individuals with stage I-IVb human papillomavirus (HPV) negative and HPV positive head and neck cancer were analyzed. The total cohort included 325 individual patients and contained 1,155 samples. Patients diagnosed with Epstein-Barr virus-associated nasopharyngeal cancer were excluded. Of the total cohort, 1,119 out of 1,155 samples (96.9%) met the cfDNA quantity quality threshold and underwent cfMeDIP processing. Samples were split into separate sets to train and test machine learning classifiers with differential methylation regions. As shown in Figure 12, blood sampling points included at diagnosis, before curative treatment (baseline (BL)), and approximately 3 months (landmark point, B1), 12 months (B2), and 24 months (B3) after curative treatment. Curative treatment included surgery alone, radiotherapy (RT) + / - chemotherapy, or surgery + radiotherapy + / - chemotherapy. 5–10 ng of plasma cfDNA was used for each sample and subjected to a bisulfite-free, non-degradable, genome-wide methylome enrichment platform based on cfMeDIP-seq, as further detailed in Example 1, Figure 25 (Workflow 4). MRD signals were quantified from mean normalized counts across informative methylated regions and binarized into positive and negative groups. Relapse-free survival (RFS) was compared between test-positive and test-negative patients 3 months after curative treatment and longitudinally.
[0179]
[0202] A total of 173 samples from 52 individual patients (Stage I (33%), Stage II (17%), Stage III (23%), Stage IV (27%)) were analyzed and correlated with relapse in this provisional training outcome. At landmark time points, patients who were test-positive showed significantly worse RFS than those who were test-negative (hazard ratio (HR) 8.91; 95% CI, 3.14–25.26, P<0.001). When continuous longitudinal samples were included, RFS continued to be statistically significantly worse in test-positive patients than in test-negative patients (HR 10.47; 95% CI, 3.81–28.82, P<0.001).
[0180]
[0203] Similarly, a larger total sample size of 249 samples collected from 75 patients was also analyzed. The clinicopathological information for the 75 patients is shown in Table 9. The analysis involved comparing the prognosis of cancer patients with and without detected ctDNA. RFS between the two groups was compared using a two-sided log-rank test with the Kaplan-Meier (KM) method. Intergroup HRs were estimated using a Cox proportional hazards model. These analyses were performed at landmark time points (Figure 13A) and longitudinally (Figure 13B). As shown in Figures 13A and 13B, ctDNA positivity was found to predict survival outcomes in patients with head and neck cancer who received curative treatment. ctDNA positivity correlated with RFS at landmark time points and longitudinally. Significant differences in RFS were observed at landmark time points with an HR of 10.97 (CI: 4.76–25.29; p<0.001) and longitudinally with an HR of 22.83 (CI: 2.8–186; p<0.001) when patients were stratified by ctDNA status. These preliminary training analyses demonstrated that MRD detection using a blood-based, tissue-independent, genome-wide methylome enrichment platform correlated strongly with RFS in HNC patients after curative treatment, with hazard ratios comparable to those of tumor-based assays.
[0181]
[0204] Furthermore, a genome-wide methylome enrichment platform based on cfMEDIP-seq was found to be capable of monitoring ctDNA dynamics, which are quantitative changes in ctDNA levels over time. To demonstrate the ability to monitor ctDNA dynamics, ctDNA was quantified and plotted over time for all relapsed and non-relapsed subjects. As shown in Figure 14, the ctDNA quantitative trajectory plots showed correlation with non-relapse and relapse outcomes. Estimated ctDNA quantification values from pre-treatment (BL) to post-treatment (e.g., B1, B2, B3) were consistent with expected ctDNA dynamics. Individual patient ctDNA before and after curative treatment was also analyzed. Figure 15 shows a representative case study of ctDNA dynamics in three individual patients (Patient A, Patient B, Patient C). Patient A (Figure 15, upper left) was diagnosed with a stage III HPV-negative tumor in the oropharynx. ctDNA was detected at diagnosis before RT treatment, and at 315 and 413 days post-diagnosis, with distant recurrence confirmed at 502 days (lead time 187 days). Patient B (Figure 15, upper right) was diagnosed with a stage IVA HPV-positive tumor in the hypopharynx. ctDNA was detected at diagnosis prior to chemoradiotherapy (CRT) treatment, but not at 107 days post-diagnosis or after completion of CRT. ctDNA was detected again 413 days after diagnosis, and distant recurrence was confirmed at 458 days (lead time 45 days). Patient C (Figure 15, bottom) presented with a stage I HPV-positive tumor in the oropharynx. ctDNA was detected at diagnosis prior to radiotherapy (RT), but not at any point after treatment. Patient C remained disease-free until the final clinical follow-up (>2.5 years). These data demonstrate that a blood-based, tissue-independent, genome-wide methylome enrichment platform can demonstrate robust performance for ctDNA quantification and monitoring of ctDNA dynamics in HPV-positive and HPV-negative head and neck cancers.
[0182]
[0205] In summary, these results can be helpful in determining whether to escalate or de-escalate treatment after curative therapy and can be used to monitor for recurrence before clinical or radiographic findings appear.
[0183] [Table 9]
[0184] Example 5: Pre-analytical variables and quality control for robust processing of plasma-derived cell-free DNA (cfDNA) using a total methylome enrichment platform.
[0206] Cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) was developed as a non-degradable liquid biopsy approach to circumvent the limitations of bisulfite sequencing. However, bisulfite-free approaches may require consistently high methylation-binding specificity to detect potentially rare methylation events in circulating tumor DNA (ctDNA). We improved cfMeDIP-seq with the aim of developing a robust genome-wide methylome enrichment platform for clinical use. This experiment investigated the effects of pre-analysis variables and analytical configuration on the platform's methylation specificity (e.g., pull-down methylated region specificity).
[0185]
[0207] Plasma-derived cfDNA was used in a genome-wide methylome enrichment platform. The cfDNA was subjected to standard library preparation, combined with DNA fillers, denatured, and subjected to immunoprecipitation using an anti-5-mC antibody. Methylation-enriched captured DNA was amplified and sequenced as described in Example 1, Figure 25 (Workflow 4). Methylation specificity of the immunoprecipitation step (e.g., pull-down methylation region specificity) was monitored using methylated and unmethylated spike-in DNA fragments added before adapter ligation. Methylation specificity was evaluated for several pre-analytical variables: sample collection tube type (Streck, EDTA), sample storage period (0–5 years, 5–10 years, ≥10 years), and genomic DNA (gDNA) contamination (1%–50% gDNA). Each variable was evaluated as the difference in mean methylation specificity (categorical variable) or correlation (continuous variable) in a cohort of over 4,000 stored plasma samples from individuals with and without cancer. The immunoprecipitation process was further optimized using 20 replicate experiments obtained from donor-derived cfDNA to enhance the centrality of methylation specificity and reduce variability.
[0186]
[0208] Methylation specificity (e.g., specificity of pulled-down methylated regions) was consistent regardless of tube type (Δ=0.29%) or sample storage period (Δ<0.32%). There was no correlation between methylation specificity when comparing samples with gDNA contamination ranging from 1% to 50% (Kendall correlation = 0.0064). Modifications to the immunoprecipitation process to improve methylation specificity were evaluated, and significant improvements were observed after optimization (Kolmogorov-Smirnov test p-value < 0.00001). The mean methylation specificity of the optimized assay was 99.7%, with 20 out of 20 samples exhibiting specificity of 99.6% or higher.
[0187]
[0209] The data demonstrated that the genome-wide methylome enrichment platform is a robust and versatile test with minimal influence from pre-analysis variables and consistently high methylation specificity.
[0188] Example 6: Analytical performance of a genome-wide methylome enrichment platform for detecting minimal residual disease from plasma-derived cell-free DNA
[0210] Tissue-independent approaches to detecting cancer signals from plasma can offer significant advantages, particularly when tissue evaluation is difficult or when rapid responses are required. Plasma-derived cell-free DNA (cfDNA) can be used to detect cancer, including minimal residual disease (MRD), in patients who have received curative cancer treatment. However, these tests may require highly sensitive cancer signal detection methods for clinical application. Therefore, genome-wide methylome enrichment platforms using cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) can differentiate cancer signals from non-cancer signals when combined with custom algorithms that leverage differential methylation regions (DMRs) found in cfDNA. Accordingly, we investigated analytical performance metrics for algorithms under development to detect MRD using methylation-based approaches.
[0189]
[0211] Plasma-derived cfDNA was used in a genome-wide methylome enrichment platform. As shown in Figure 16 and further detailed in Example 1 (Workflow 4 in Figure 25), plasma-derived cfDNA (10 ng input) combined with spike-in DNA fragments was subjected to standard library preparation, combined with DNA fillers, denatured, and subjected to immunoprecipitation using an anti-5-mC antibody. The methylation-enriched captured DNA was amplified and sequenced. The sequencing results were subjected to bioinformatics pipeline and algorithmic classification. Cancer-specific methylation was quantified using candidate algorithms consisting of DMRs. The detection limits and precision were then evaluated using artificial cancer samples intended to mimic low levels of circulating tumor DNA (ctDNA), represented by MRD.
[0190]
[0212] To evaluate the detection limit, samples were prepared by enzymatically fragmenting DNA from three immortalized tumor-derived cell lines (H-1299 lung cancer cell line, FaDu head and neck (H&N) cancer cell line, and A-253 H&N cancer cell line) and titrating it to pooled cfDNA from non-cancer donors using a titration series targeting less than 1% ctDNA levels. As outlined in Figure 17, five technical iterations were used for each ctDNA level, resulting in a total of 65 individual cfMeDIP runs. As shown in Figure 18, all non-cancer and induced cancer samples met the in-process quality control criteria, including methylation specificity of over 98.5% and more than 80 million unique molecules. Furthermore, as shown in Figure 19, ctDNA methylation scores were measured for different levels of ctDNA, and the limit of detection (LoD95) calculation at 95% sensitivity was found to be less than 0.1% using a tissue-independent algorithm. The threshold was set using 12 non-cancer donors and 5 replicates of the non-cancer pool, establishing a true negative rate of 95%.
[0191]
[0213] To evaluate the accuracy of the blood-based, tissue-independent, genome-wide methylome enrichment platform, samples were prepared in 18 technical iterations, as outlined in Figure 20, by enzymatically fragmenting DNA from two immortalized tumor cell lines (A-253 H&N cancer cell line and FaDu H&N cancer cell line) at high (1.12%), medium (0.84%), and low (0.56%) ctDNA levels, titrating it to pooled non-cancer donor-derived cfDNA. The samples were then processed using the blood-based, tissue-independent, genome-wide methylome enrichment platform in repeated tests with various operators, sequencing runs, and antibody reagent lots. As shown in Figure 21, the ctDNA methylation score was used to assess agreement with expected results and variability. At all levels of ctDNA, the results were consistent with expected results. Furthermore, as shown in Figure 22, the combination of all-variable analyses (operator, sequencing run, and antibody reagent lot) showed a dispersion component (CV%) of less than 40% at all ctDNA levels.
[0192]
[0214] In summary, the evaluation of detection limits and accuracy demonstrated the use of a blood-based, tissue-independent, genome-wide methylome enrichment platform utilizing a non-degradable method combining specific algorithms with DMR, at a level suitable for MRD detection.
[0193] Example 7: Prognostic performance of a genome-wide methylome enrichment platform in early non-small cell lung cancer (NSCLC)
[0215] Circulating tumor DNA (ctDNA) can be used to identify the presence of cancer and minimal residual disease (MRD). Quantifying ctDNA may be a useful cancer management tool for assessing prognosis. In this example, we evaluated the potential of quantifying ctDNA in plasma using a tumor-independent, genome-wide methylome enrichment platform to predict recurrence in early non-small cell lung cancer (NSCLC).
[0194]
[0216] Retrospective evaluation was performed using preserved pre-treatment samples (collected 2009–2013, Princess Margaret Cancer Centre) from newly diagnosed stage I and II NSCLC patients. Blood samples were obtained after cancer diagnosis and before treatment. As described in Example 1, Figure 25 (Workflow 4), samples were analyzed using 5–10 ng of cell-free DNA isolated from plasma on a bisulfite-free, non-degradable, genome-wide DNA methylation enrichment platform. Samples from 41 patients were included. Table 10 lists the clinical and demographic information of the patients. ctDNA was quantified from the mean normalized count across the entire information region. An event was defined as cancer recurrence or death from any cause, whichever came first. The ctDNA amount threshold was set at the point where 100% of samples from patients without an event were below the threshold (i.e., 100% specificity). For samples with ctDNA amounts above the threshold, the time to recurrence or death was compared to ctDNA amounts below the threshold. Relapse-free survival was estimated using the Kaplan-Meier method, comparing samples with ctDNA levels above the threshold with those below the threshold. The difference between the two groups was evaluated by the log-rank test. Multivariate Cox regression analysis was used to adjust for known prognostic values of pretreatment ctDNA levels and for covariates identified in univariate analysis.
[0195]
[0217] Twenty-seven events occurred, and the median follow-up period was 55.8 months. Samples with ctDNA above the threshold showed significantly worse recurrence-free survival, as shown in Figure 11 [hazard ratio (HR) 2.70 (95% CI 1.26, 5.78), log-rank P=0.008]. Multivariate analysis (Table 11) showed that even after accounting for histology (selected using univariate analysis), samples with ctDNA above the threshold showed significantly worse recurrence-free survival [HR 2.79 (95% CI 1.30, 6.02), P=0.009].
[0196]
[0218] In summary, the data demonstrated that this blood-based, genome-wide methylome enrichment platform can be used for ctDNA quantification and prognosis prediction in early NSCLC. Here, initial feasibility was assessed using plasma samples from treatment-naive patients. Further evaluation of its application to cancer management will be needed in future studies utilizing post-treatment and longitudinal samples.
[0197] [Table 10]
[0198] [Table 11]
[0199] Example 8: Cancer methylome vs. total methylome
[0219] Various amounts of nucleosomal cell-free DNA (ncfDNA) derived from the FaDu cancer cell line were generated to mimic cancer signals and diluted against a background of pooled cfDNA from multiple cancer-free donors. Table 12 shows the dilution points and technical iterations for a total of 32 samples prepared for use in limit of detection (LoD) evaluation.
[0200] [Table 12]
[0201]
[0220] Ten ng of input DNA from each of the 32 samples was subjected to a genome-wide DNA methylation enrichment platform based on cell-free methylated DNA immunoprecipitation to generate enriched libraries, as described in Example 1, Figure 25 (Workflow 4). After cell-free methylated DNA immunoprecipitation, the enriched libraries were mixed with probes targeting the desired target regions, as shown in Figure 28. Two target capture pools were prepared by equally dividing the enriched libraries generated from each dilution (Table 13), ensuring that the technical replicates for each dilution were equal within each pool, resulting in a total of 16 samples per pool. Each pool underwent a capture process using a set of probes targeting the following cancer methylome target panel, followed by the Twist Target Enrichment Standard Hybridization v2 protocol to generate enriched libraries for sequencing on an Illumina next-generation sequencing system. The cancer methylome target panel consisted of the following: 1) housekeeping methylation regions that are methylated across different samples regardless of cancer status; 2) CpG-free regions used to calculate the endogenous binding specificity score for quantifying nonspecific methylation bonds as a quality control measure; 3) highly differential methylation regions (DMRs) of cancer; 4) noise regions; and 5) regions with low counts in cancer-free samples (controls).
[0202] [Table 13]
[0203]
[0221] Target enrichment samples were sequenced using an Illumina NovaSeq 6000 sequencer, targeting approximately 26 million reads per sample. Next, sequenced reads from the target capture samples were aligned to the hg38 human reference genome using the Bowtie2 alignment tool. Sequence duplications potentially caused by PCR or flow-cell optical clustering errors were identified and removed using a molecular barcode (UMI)-based approach. DNA sequences were further filtered by size, minimum number of CpGs, and mapping quality before being counted in the region of interest (ROI) of the target genome.
[0204]
[0222] Cancer DNA was quantified by generating a score by counting the number of DNA fragments in the target ROI and comparing it to an internal baseline. Using the scores obtained from the same ROI, a cancer methylome approach that integrated an additional enrichment step using probes was compared with a whole methylome approach that did not integrate additional target region capture using probes. The whole methylome results were based on previous cell line titration studies and the highest performance score. The same region was used to obtain the cancer methylome score. Next, the change factor was calculated by dividing the LoD of the whole methylome approach by the LoD of the cancer methylome approach. The LoD was determined in both approaches using ctDNA quantification values across different dilution points. As shown in Figure 27, there was an improvement in LoD in the cancer methylome approach compared to the whole methylome approach, as shown by the change factor, indicating that the LoD can be improved by integrating an additional enrichment step using probes.
[0205]
[0223] Preferred embodiments of the present invention have been described herein, but it will be understood by those skilled in the art that modifications can be made therewith without departing from the spirit of the invention or the scope of the appended claims. All documents disclosed herein, including those in the following list of references, are incorporated by reference.
Claims
1. (a) A step of obtaining a first nucleic acid molecule from the target cell-free sample, (b) A step of producing a second set of nucleic acid molecules from a first set of nucleic acid molecules or derivatives thereof, wherein the methylation level of the second set of nucleic acid molecules is concentrated compared to that of the first set of nucleic acid molecules, (c) A step of concentrating one or more targets in the second set of nucleic acid molecules or derivatives thereof to obtain a third set of nucleic acid molecules, (d) A step of sequencing the nucleic acid molecules or derivatives of the third set; A method that includes this.
2. The method according to claim 1, wherein the concentration step includes a step of contacting the second set of nucleic acid molecules or derivatives thereof with one or more nucleic acid capture probes.
3. The method according to claim 1 or 2, wherein the generating step includes contacting the nucleic acid molecules or derivatives of the first set with a methylated nucleic acid scavenging reagent.
4. The method according to claim 3, wherein the methylated nucleic acid scavenging reagent is formed by incubating a methylated bond molecule with a solid substrate (ii).
5. The method according to claim 4, wherein the solid substrate is a bead.
6. The method according to claim 4 or 5, wherein the solid substrate is a magnetic solid substrate.
7. The method according to any one of claims 4 to 6, wherein the solid substrate contains protein A.
8. The method according to any one of claims 4 to 6, wherein the solid substrate contains streptavidin.
9. The method according to any one of claims 4 to 8, wherein the methylated bond molecule is an antibody.
10. The method according to any one of claims 4 to 9, wherein the methylated bond molecule comprises biotin.
11. The method according to any one of claims 4 to 10, wherein the methylated bond molecule is bonded to methylated cytosine.
12. The method according to any one of claims 4 to 11, further comprising the step of amplifying the second set of molecules.
13. The method according to claim 12, wherein the amplification step is performed while a subset of the first set of nucleic acids is bound to the methylated nucleic acid capture reagent.
14. The method according to any one of claims 1 to 13, wherein the sequencing step is sequencing by a synthesis reaction.
15. The method according to any one of claims 1 to 14, wherein the sequencing step does not include bisulfite sequencing.
16. The method according to any one of claims 1 to 15, wherein one or more library preparation reactions are performed on the third set of nucleic acid molecules prior to the sequencing step.
17. The method according to claim 16, further comprising the step of incubating the third set of nucleic acid molecules with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out the one or more library preparation reactions and before the sequencing step.
18. The method according to claim 17, further comprising the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, and then subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with the nucleic acid.
19. The method according to claim 18, further comprising the step of performing additional magnetic trapping.
20. The method according to any one of claims 1 to 19, further comprising the step of adding a certain amount of filler DNA to the first nucleic acid molecule before (b).
21. (a) A step of providing multiple nucleic acid molecules generated from a target cell-free deoxyribonucleotide (cfDNA) sample, (b) A step of subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads, (c) A step of computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules, (d) A step of computer processing the methylation profile to determine that the subject has cancer in an area under the receiver operating characteristic curve (AUROC) of at least about 91%, wherein the cancer is low-exudative cancer. A method that includes this.
22. The method according to claim 21, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from a low-exudative tumor.
23. The method according to claim 21, wherein the cancer is bladder cancer, breast cancer, uterine cancer, prostate cancer, or kidney cancer.
24. The method according to claim 23, wherein the cancer is endometrial cancer or prostate cancer.
25. (a) A step of providing multiple nucleic acid molecules generated from a target cell-free deoxyribonucleotide (cfDNA) sample, (b) A step of subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads, (c) A step of computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules, (d) A step of computer processing the methylation profile to determine that the subject has cancer in an area under the receiver operating characteristic curve (AUROC) of at least about 94%, wherein the cancer is an early-stage cancer. A method that includes this.
26. The method according to claim 25, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from an early tumor.
27. The method according to claim 26, wherein the early tumor is a stage I tumor.
28. The method according to claim 26, wherein the early tumor is a stage II tumor.
29. The method according to any one of claims 25 to 28, wherein the cancer is bladder cancer, breast cancer, colorectal cancer, uterine cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, prostate cancer, or kidney cancer.
30. The method according to claim 29, wherein the cancer is esophageal cancer, hepatobiliary tract cancer, or ovarian cancer.
31. (a) A step of providing multiple nucleic acid molecules generated from a target cell-free deoxyribonucleotide (cfDNA) sample, (b) A step of subjecting the plurality of nucleic acid molecules or derivatives thereof to sequencing to generate a plurality of sequencing reads in the absence of bisulfite conversion, (c) A step of computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules, (d) A step of computer processing the methylation profile to determine whether the subject has cancer, wherein the cancer is endometrial cancer, esophageal cancer, hepatobiliary cancer, ovarian cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer. A method that includes this.
32. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from endometrial cancer.
33. The method according to claim 32, wherein the methylation profile is computer-processed to determine that the subject has endometrial cancer in at least about 90% of the area under the receiver operating characteristic curve (AUROC).
34. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from esophageal cancer.
35. The method according to claim 34, wherein the methylation profile is computer-processed to determine that the subject has esophageal cancer in at least about 99% of the area under the receiver operating characteristic curve (AUROC).
36. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from hepatobiliary tract cancer.
37. The method according to claim 36, wherein the methylation profile is computer-processed to determine that the subject has hepatobiliary tract cancer with at least about 99% of the area under the receiver operating characteristic curve (AUROC).
38. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from ovarian cancer.
39. The method according to claim 38, wherein the methylation profile is computer-processed to determine that the subject has ovarian cancer in at least about 97% of the area under the receiver operating characteristic curve (AUROC).
40. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from prostate cancer.
41. The method according to claim 40, wherein the methylation profile is computer-processed to determine that the subject has prostate cancer in at least about 89% of the area under the receiver operating characteristic curve (AUROC).
42. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from bladder cancer.
43. The method according to claim 42, wherein the methylation profile is computer-processed to determine that the subject has bladder cancer in at least about 95% of the area under the receiver operating characteristic curve (AUROC).
44. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from breast cancer.
45. The method according to claim 44, wherein the methylation profile is computer-processed to determine that the subject has breast cancer in at least about 92% of the area under the receiver operating characteristic curve (AUROC).
46. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from colorectal cancer.
47. The method according to claim 46, wherein the methylation profile is computer-processed to determine that the subject has colorectal cancer in at least about 98% of the area under the receiver operating characteristic curve (AUROC).
48. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from head and neck cancer.
49. The method according to claim 48, wherein the methylation profile is computer-processed to determine that the subject has head and neck cancer in at least about 96% of the area under the receiver operating characteristic curve (AUROC).
50. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from lung cancer.
51. The method according to claim 50, wherein the methylation profile is computer-processed to determine that the subject has lung cancer in at least about 96% of the area under the receiver operating characteristic curve (AUROC).
52. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from pancreatic cancer.
53. The method according to claim 52, wherein the methylation profile is computer-processed to determine that the subject has pancreatic cancer in at least about 99% of the area under the receiver operating characteristic curve (AUROC).
54. The method according to claim 31, wherein the cfDNA sample contains circulating tumor nucleic acid molecules derived from kidney cancer.
55. The method according to claim 54, wherein the methylation profile is computer-processed to determine that the subject has renal cancer in at least about 91% of the area under the receiver operating characteristic curve (AUROC).
56. The method according to any one of claims 21 to 55, further comprising the step of adding a set of nucleic acid molecules not derived from the subject to the plurality of nucleic acid molecules before (b).
57. The method according to any one of claims 21 to 56, wherein the methylation profile is genome-wide.
58. The method according to claim 57, wherein the methylation profile comprises all methylomes.
59. The method according to any one of claims 21 to 58, wherein (d) includes a supervised machine learning method, the supervised machine learning method being regression, support vector machine, tree-based method, neural network, or nearest neighbor method.
60. The method according to any one of claims 21 to 58, wherein (d) includes an unsupervised machine learning method, the unsupervised machine learning method being clustering, a neural network, principal component analysis, or matrix factorization.
61. The method according to any one of claims 21 to 60, wherein the subject has previously received treatment for cancer, the cancer has substantially disappeared, and (d) determines that the subject has a recurrence of the cancer.
62. The method according to any one of claims 21 to 61, further comprising the step of adding a certain amount of filler DNA to the plurality of nucleic acid molecules or derivatives before (b).
63. The method according to claim 62, wherein the filler DNA includes double-stranded DNA.
64. The method according to any one of claims 62 to 63, wherein the amount of filler DNA is about 20 nanograms (ng) to about 100 ng.
65. The method according to any one of claims 62 to 64, wherein at least a portion of the filler DNA is methylated.
66. The method according to claim 65, wherein 10% to 40% of the filler DNA is methylated and the remainder is unmethylated filler DNA.
67. The method according to any one of claims 21 to 66, further comprising the step of contacting the cfDNA sample with a methylated nucleic acid capture reagent before (b) to generate the plurality of nucleic acid molecules, wherein the plurality of nucleic acids include one or more methylated regions.
68. The method according to claim 67, wherein the methylated nucleic acid scavenger comprises a binder and a solid substrate.
69. The method according to claim 68, wherein the methylated nucleic acid capture reagent is produced by coupling the binder to the solid substrate by incubating the binder together with the solid substrate.
70. The method according to claim 69, wherein coupling the binder to the solid substrate is performed before the step of contacting the cfDNA sample with the methylated nucleic acid capture reagent.
71. The method according to any one of claims 68 to 70, wherein the solid substrate is beads.
72. The method according to any one of claims 68 to 70, wherein the solid substrate is protein A beads.
73. The method according to any one of claims 68 to 72, wherein the solid substrate is a magnetic solid substrate.
74. The method according to any one of claims 68 to 73, wherein the binder comprises an antibody.
75. The method according to any one of claims 68 to 74, wherein the binder is selected from the group consisting of an anti-5-methylcytosine antibody or a derivative thereof, an anti-5-carboxylcytosine antibody or a derivative thereof, an anti-5-formylcytosine antibody or a derivative thereof, an anti-5-hydroxymethylcytosine antibody or a derivative thereof, an anti-3-methylcytosine antibody or a derivative thereof, and any combination thereof.
76. The method according to any one of claims 67 to 75, wherein one or more methylated regions are enriched in the plurality of nucleic acids with a specificity of at least about 99%.
77. The method according to any one of claims 21 to 76, further comprising the step of amplifying the plurality of nucleic acid molecules to produce an amplicon before (b), wherein (c) comprises the step of sequencing the amplicon.
78. The method according to claim 77, wherein the amplification step is performed on the plurality of nucleic acids while the plurality of nucleic acids are bound to a solid support.
79. The method according to any one of claims 77 to 78, wherein the amplification step includes PCR amplification.
80. The method according to claim 79, wherein the PCR amplification comprises at least 13 cycles or at least 14 cycles.
81. The method according to any one of claims 21 to 80, further comprising the step of contacting the plurality of nucleic acid molecules with one or more nucleic acid capture probes to concentrate one or more target sequences before (b).
82. The method according to claim 81, wherein the one or more target sequences include one or more genes.
83. A method for processing nucleic acid samples from a target, (a) A step of producing a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the subject, wherein the plurality of filler nucleic acids comprises at least one methylated nucleic acid molecule, (b) A step of incubating (i) a methylated bond molecule with (ii) a solid substrate to form a methylated nucleic acid scavenging reagent, (c) A step of capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, thereby concentrating the plurality of methylated nucleic acids in the nucleic acid sample mixture. A method that includes this.
84. The method according to claim 83, further comprising the step of amplifying the captured methylated nucleic acid after the capturing step to generate an amplicon of the plurality of methylated nucleic acids.
85. The method according to claim 84, wherein the amplification step is performed while the plurality of methylated nucleic acids are bound to the methylated nucleic acid scavenging reagent.
86. The method according to claim 84 or 85, wherein the captured methylated nucleic acid is not subjected to an elution reaction before amplification.
87. The method according to any one of claims 83 to 86, further comprising the step of subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction.
88. The method according to claim 87, wherein the sequencing reaction is sequencing by a synthesis reaction.
89. The method according to claim 87 or 88, wherein the sequencing reaction does not include bisulfite sequencing.
90. The method according to any one of claims 83 to 89, wherein the plurality of methylated nucleic acids include cell-free nucleic acids.
91. The method according to any one of claims 83 to 90, wherein the solid substrate is beads.
92. The method according to any one of claims 83 to 91, wherein the solid substrate is a magnetic solid substrate.
93. The method according to any one of claims 83 to 92, wherein the solid substrate contains protein A.
94. The method according to any one of claims 83 to 93, wherein the solid substrate contains streptavidin.
95. The method according to any one of claims 83 to 94, wherein the methylated bond molecule is an antibody.
96. The method according to any one of claims 83 to 95, wherein the methylated bond molecule comprises biotin.
97. The method according to any one of claims 83 to 96, wherein the methylated bond molecule is bonded to methylated cytosine.
98. The method according to any one of claims 83 to 97, further comprising the steps of obtaining the nucleic acid sample from the subject and carrying out one or more library preparation reactions on the nucleic acid sample, prior to (a).
99. The method according to claim 98, further comprising the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out one or more library preparation reactions and before (a).
100. The method according to claim 99, further comprising the steps of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, and then subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids.
101. The method according to claim 100, further comprising the step of performing additional magnetic trapping.
102. The method according to any one of claims 83 to 101, further comprising the step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (c).
103. The method according to any one of claims 83 to 102, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated by at least twofold.
104. The method according to any one of claims 83 to 103, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 100 times.
105. The method according to any one of claims 83 to 104, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99% specificity.
106. The method according to any one of claims 83 to 105, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99.5% specificity.
107. The method according to any one of claims 83 to 106, further comprising the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
108. The method according to claim 107, wherein the one or more target sequences include one or more genes.
109. A method for processing nucleic acid samples from a target, (a) A step of providing a nucleic acid sample containing multiple methylated nucleic acids, (b) A step of incubating (i) a methylated bond molecule with (ii) a solid substrate to form a methylated nucleic acid scavenging reagent, (c) A step of capturing the plurality of methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, thereby concentrating the plurality of methylated nucleic acids in the nucleic acid sample mixture, wherein the plurality of methylated nucleic acids are concentrated with a specificity higher than 99%. A method that includes this.
110. The method according to claim 109, further comprising the step of amplifying the captured methylated nucleic acid after the capturing step to generate amplicons of a plurality of methylated nucleic acids.
111. The method according to claim 110, wherein the amplification step is performed while the plurality of methylated nucleic acids are bound to the methylated nucleic acid scavenging reagent.
112. The method according to claim 109 or 110, wherein the captured methylated nucleic acid is not subjected to an elution reaction before amplification.
113. The method according to any one of claims 109 or 112, further comprising the step of subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction.
114. The method according to claim 113, wherein the sequencing reaction is sequencing by a synthesis reaction.
115. The method according to claim 113 or 114, wherein the sequencing reaction does not include bisulfite sequencing.
116. The method according to any one of claims 109 to 115, wherein the plurality of methylated nucleic acids include cell-free nucleic acids.
117. The method according to any one of claims 109 to 116, wherein the solid substrate is a bead.
118. The method according to any one of claims 109 to 117, wherein the solid substrate is a magnetic solid substrate.
119. The method according to any one of claims 109 to 118, wherein the solid substrate contains protein A.
120. The method according to any one of claims 109 to 119, wherein the solid substrate contains streptavidin.
121. The method according to any one of claims 109 to 120, wherein the methylated molecule is an antibody.
122. The method according to any one of claims 109 to 121, wherein the methylated bond molecule comprises biotin.
123. The method according to any one of claims 109 to 122, wherein the methylated bond molecule is bonded to methylated cytosine.
124. The method according to any one of claims 109 to 123, further comprising the step of carrying out one or more library preparation reactions on the methylated nucleic acid before (c).
125. The method according to claim 124, further comprising the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out one or more library preparation reactions and before (c).
126. The method according to claim 125, further comprising the steps of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, and then subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids.
127. The method according to claim 126, further comprising the step of performing additional magnetic trapping.
128. The method according to any one of claims 109 to 127, further comprising the step of denaturing the nucleic acid in the nucleic acid sample after (a) and before (c).
129. The method according to any one of claims 109 to 128, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated by at least twofold.
130. The method according to any one of claims 109 to 129, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated at least 100 times.
131. The method according to any one of claims 109 to 130, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated with a specificity of at least 99.5%.
132. The method according to any one of claims 109 to 131, further comprising the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
133. The method according to claim 132, wherein the one or more target sequences include one or more genes.
134. A method for processing nucleic acid samples from a target, (a) A step of providing a nucleic acid sample containing multiple methylated nucleic acids, (b) A step of incubating (i) a methylated bond molecule with (ii) a solid substrate to form a methylated nucleic acid scavenging reagent, (c) A step of capturing the methylated nucleic acid by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating a methylated nucleic acid to which a solid substrate is bound; (d) A step of amplifying the methylated nucleic acid to which the solid substrate is bound to generate an amplicon of the plurality of methylated nucleic acids. A method that includes this.
135. The method according to claim 134, wherein the amplification step is performed while the plurality of methylated nucleic acids are bound to the methylated nucleic acid scavenging reagent.
136. The method according to claim 134 or 135, wherein the captured methylated nucleic acid is not subjected to an elution reaction before amplification.
137. The method according to any one of claims 134 to 136, further comprising the step of subjecting the amplicon of a methylated nucleic acid to a sequencing reaction.
138. The method according to claim 137, wherein the sequencing reaction is sequencing by a synthesis reaction.
139. The method according to claim 137 or 138, wherein the sequencing reaction does not include bisulfite sequencing.
140. The method according to any one of claims 134 to 139, wherein the plurality of methylated nucleic acids include cell-free nucleic acids.
141. The method according to any one of claims 134 to 140, wherein the solid substrate is a bead.
142. The method according to any one of claims 134 to 141, wherein the solid substrate is a magnetic solid substrate.
143. The method according to any one of claims 134 to 142, wherein the solid substrate contains protein A.
144. The method according to any one of claims 134 to 143, wherein the solid substrate contains streptavidin.
145. The method according to any one of claims 134 to 144, wherein the methylated bond molecule is an antibody.
146. The method according to any one of claims 134 to 145, wherein the methylated bond molecule comprises biotin.
147. The method according to any one of claims 134 to 146, wherein the methylated bond molecule is bonded to methylated cytosine.
148. The method according to any one of claims 134 to 147, wherein one or more library preparation reactions are carried out on the methylated nucleic acid before (c).
149. The method according to claim 148, further comprising the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out one or more library preparation reactions and before (c).
150. The method according to claim 149, further comprising the steps of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, and then subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids.
151. The method according to claim 150, further comprising the step of performing additional magnetic trapping.
152. The method according to any one of claims 134 to 151, further comprising the step of denaturing the nucleic acid in the nucleic acid sample after (a) and before (c).
153. The method according to any one of claims 134 to 152, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated by at least twofold.
154. The method according to any one of claims 134 to 153, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated at least 100 times.
155. The method according to any one of claims 134 to 154, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated with at least 99% specificity.
156. The method according to any one of claims 134 to 155, wherein the plurality of methylated nucleic acids in the nucleic acid sample are concentrated with at least 99.5% specificity.
157. The method according to any one of claims 134 to 156, further comprising the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
158. The method according to claim 157, wherein the one or more target sequences include one or more genes.
159. A method for processing nucleic acid samples from a target, (a) A step of producing a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the subject, wherein the plurality of filler nucleic acids comprises at least one methylated nucleic acid molecule, (b) A step of capturing the methylated nucleic acid by adding a capture reagent containing a solid substrate to the nucleic acid sample mixture, thereby generating methylated nucleic acid to which the solid substrate is bound, (c) A step of amplifying the methylated nucleic acid to which the solid substrate is bound to generate an amplicon of the methylated nucleic acid. A method that includes this.
160. The method according to claim 159, wherein the captured methylated nucleic acid is not subjected to an elution reaction before amplification.
161. The method according to claim 159 or 160, further comprising the step of subjecting the amplicon of the methylated nucleic acid to a sequencing reaction.
162. The method according to claim 161, wherein the sequencing reaction is sequencing by a synthesis reaction.
163. The method according to claim 161 or 162, wherein the sequencing reaction does not include bisulfite sequencing.
164. The method according to any one of claims 159 to 163, wherein the plurality of methylated nucleic acids include cell-free nucleic acids.
165. The method according to any one of claims 159 to 164, wherein the solid substrate is a bead.
166. The method according to any one of claims 159 to 165, wherein the solid substrate is a magnetic solid substrate.
167. The method according to any one of claims 159 to 166, wherein the solid substrate contains protein A.
168. The method according to any one of claims 159 to 167, wherein the solid substrate contains streptavidin.
169. The method according to any one of claims 159 to 168, wherein the capture reagent is a methylated nucleic acid capture agent.
170. The method according to claim 169, wherein the methylated nucleic acid scavenging agent includes methylated bond molecules attached to the solid substrate.
171. The method according to claim 170, wherein the methylated bond molecule is an antibody.
172. The method according to claim 170, wherein the methylated bond molecule contains biotin.
173. The method according to claim 172, wherein the methylated bond molecule is bonded to methylated cytosine.
174. The method according to any one of claims 159 to 173, further comprising the steps of obtaining the nucleic acid sample from the subject and carrying out one or more library preparation reactions on the nucleic acid, prior to (a).
175. The method according to claim 174, further comprising the step of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, after the step of carrying out one or more library preparation reactions and before (a).
176. The method according to claim 175, further comprising the steps of incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, and then subjecting the sample to magnetic capture to remove the plurality of magnetic beads that interact with nucleic acids.
177. The method according to claim 176, further comprising the step of performing additional magnetic trapping.
178. The method according to any one of claims 159 to 177, further comprising the step of denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b).
179. The method according to any one of claims 159 to 178, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated by at least twofold.
180. The method according to any one of claims 159 to 179, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated at least 100 times.
181. The method according to any one of claims 159 to 180, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with at least 99% specificity.
182. The method according to any one of claims 159 to 181, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are concentrated with a specificity of at least 99.5%.
183. The method according to any one of claims 159 to 182, further comprising the step of contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
184. The method according to claim 183, wherein the one or more target sequences include one or more genes.