Methods and Systems for High-Throughput Sequencing of Methylated Nucleic Acids
The method addresses the limitations of bisulfite-based methylation sequencing by using enzymatic conversion and enrichment techniques to preserve fragment length and improve accuracy in DNA methylation analysis, enhancing sensitivity and reducing sample degradation.
Patent Information
- Application Number
- JP2021511648
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2020-05-29
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Existing methylation sequencing methods, particularly bisulfite-based approaches, cause significant degradation of cell-free DNA (cfDNA), leading to loss of fragment length information and require large sample amounts or low-depth coverage, which limits the accuracy and sensitivity of DNA methylation analysis, especially in liquid biopsy applications.
A method involving ligation of a nucleic acid adapter with a unique molecular identifier, followed by minimally disruptive enzymatic conversion of unmethylated cytosine to uracil, amplification, and enrichment using nucleic acid probes complementary to pre-identified CpG or CH loci, with deep sequencing to determine methylation profiles accurately.
This approach preserves fragment length information, improves sequencing accuracy, and reduces sample degradation, enabling high-depth genomic coverage and enhanced sensitivity in methylation analysis.
Smart Images

Figure 0007717608000003 
Figure 0007717608000004 
Figure 0007717608000005
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit of U.S. Provisional Application No. 62 / 855,795, filed May 31, 2019, which is hereby incorporated by reference in its entirety.
[0002] Incorporation by reference All publications, patents, and patent applications mentioned in this specification, including the examples, are hereby incorporated by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions in the specification, will control.
Background Art
[0003] Due to the stability and role of DNA in diseases such as normal differentiation and cancer, DNA methylation can represent tumor characteristics and phenotypic states, making it highly useful for personalized medicine. Since abnormal DNA methylation patterns occur early in cancer development, early cancer detection can be facilitated. In fact, DNA methylation abnormalities are one of the characteristics of cancer and are associated with all aspects of cancer, from tumorigenesis to cancer progression and metastasis. Due to such characteristics, in many recent techniques, DNA methylation patterns have come to be used for diagnosing cancer. Specifically, cell-free DNA (cfDNA) is fragmented DNA present in circulation, and the fragmentation pattern is useful and beneficial as a biological signal. In contrast, genomic DNA is artificially fragmented in vitro for use in library preparation, so the fragmentation pattern of genomic DNA is not very important in diagnostic methods.
[0004] DNA methylation is a covalent modification of DNA and a stable genetic feature capable of playing an important role in suppressing gene expression and regulating chromatin structure. In humans, DNA methylation mainly occurs at the cytosine residues of CpG dinucleotides. Unlike other dinucleotides, CpG is not evenly distributed throughout the genome and may be concentrated in short CpG-rich DNA regions called CpG islands. Usually, about 70-75% of most CpG sites in the genome are methylated. However, the methylation pattern varies among cell types, reflecting its role in regulating cell type-specific gene expression. In this way, the methylome of a cell can program the cell's terminal differentiation state, such as into neurons, muscle cells, immune cells, etc.
[0005] Furthermore, various cell subtypes in a tissue can exhibit different methylation patterns. In cancer cells, CpG methylation can be deregulated, and abnormal methylation patterns are part of the earliest events in tumorigenesis. The methylation profile of a given cancer type is most similar to that of the tissue from which the cancer originated. Therefore, abnormal methylation features on cfDNA fragments can be used to distinguish cancer cells from normal cells and determine the tissue type of origin. Usually, the overall CpG methylation level decreases in cancer cells, but the average methylation level (or methylation %) at specific loci may vary at specific CpG sites in cancer cells compared to paired normal cells. By profiling CpGs (DMCs, single sites) or regions (DMRs, multiple sites within a local region) with different methylation between normal and diseased cells, disease biomarkers can be identified. Such an approach led to the development of the SEPT9 gene methylation assay (Epi proColon), which is the first FDA-approved diagnostic method for colorectal cancer (CRC) using blood.
[0006] Bisulfite conversion or bisulfite sequencing has become a widely used method for DNA methylation analysis. Bisulfite sequencing is a convenient and effective method for mapping DNA methylation to individual bases. Unfortunately, bisulfite conversion is a severely destructive process for cfDNA, degrading more than 90% of the sample DNA. The two main methods for constructing bisulfite sequencing libraries are: (1) bisulfite conversion of DNA before library construction, which requires the construction of single-stranded DNA libraries, and (2) bisulfite conversion of DNA after double-stranded adapter ligation. Both involve significant degradation of DNA, which can be a problem for cfDNA, which is present in plasma at very low concentrations and is a limited resource for liquid biopsy applications. In ssDNA libraries, some degraded cfDNA can be retained within the library, but the endpoint information regarding the degraded fragments is lost. Such libraries limit the ability to use cfDNA endpoint or fragment length information to test for DNA methylation. In dsDNA libraries, cfDNA inserts cleaved by bisulfite are lost from the library, but the endpoint information regarding the surviving cfDNA inserts is retained. This requires extremely large amounts of blood sampling to achieve high-depth genomic coverage or limits the execution of the analysis to only low-depth genomic coverage.
[0007] The advent of next-generation DNA sequencing has advanced clinical medicine and basic research. However, although this technique can generate hundreds of billions of nucleotides of DNA sequence in a single experiment, the error rate is approximately 1%, resulting in hundreds of millions of sequencing errors. Such errors are acceptable for some applications but pose a major problem in "deep sequencing" of genetically heterogeneous mixtures such as tumors and mixed microbial populations.
[0008] In existing methods, two different sequencing assays and two different cfDNA pools are required to analyze cfDNA variants and the methylation status of cfDNA. From the perspective of plasma / cfDNA input and associated costs, such methods can be extremely costly. Additionally, bisulfite DNA disruption may reduce the sensitivity of variant calling methods applicable to sequencing data of bisulfite-converted DNA (compared to enzymatic conversion). Therefore, to maintain the integrity of the sample nucleic acid and improve the accuracy of methylation status analysis at the whole genome level or target level, improvement of the cfDNA methylation analysis method is required.
SUMMARY OF THE INVENTION
[0009] The methods and systems provided herein address the limitations of bisulfite-based methylation sequencing by improving the quality and accuracy of nucleic acid methylation sequencing and are used for disease detection. As the accuracy increases and the information regarding the methylation status becomes complete, the quality of feature generation used in machine learning models and classifier generation can be improved.
[0010] In a first aspect, a method for performing methylation sequencing of a nucleic acid sample is provided, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to the nucleic acid molecule, wherein the nucleic acid molecule comprises an unconverted nucleic acid; b) converting unmethylated cytosine to uracil within the nucleic acid molecule using a minimally disruptive conversion method, thereby generating a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction, thereby generating an amplified converted nucleic acid; d) probing the amplified converted nucleic acid with a nucleic acid probe complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby generating a probed converted nucleic acid. e) determining the nucleic acid sequence of the probed converted nucleic acid at a depth of more than 100x; f) comparing the nucleic acid sequence of the probed converted nucleic acid with the reference nucleic acid sequences of a pre-identified panel of CpG or CH loci to determine the methylation profile of the nucleic acid molecules in the biological sample.
[0011] In one embodiment, the nucleic acid molecule is plasma cfDNA.
[0012] In one embodiment, the minimally disruptive conversion method includes enzymatic conversion, TAPS, or CAPS.
[0013] In one embodiment, the unique molecular identifier is 4bp to 6bp in length and has a 5'thymidine overhang.
[0014] In one embodiment, the nucleic acid adapter further includes a unique dual index (UDI) sequence. In one embodiment, the length of the UDI sequence is 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 11bp, or 12bp.
[0015] In one embodiment, the step of amplifying the converted nucleic acid includes using a primer containing a unique dual index (UDI) sequence. In one embodiment, the length of the UDI sequence is 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 11bp, or 12bp.
[0016] In one embodiment, the nucleic acid adapter is a conversion-resistant adapter that includes each of the bases guanine, thymine, adenine, and cytosine but does not include 5mC-containing bases or 5hmC-containing bases.
[0017] In one embodiment, the nucleic acid probe is an unmethylated nucleic acid probe.
[0018] In one embodiment, the nucleic acid probe hybridizes to a target region of interest that matches an unmethylated cytosine at a CpG site within a reference nucleic acid sequence.
[0019] In one embodiment, the nucleic acid probe includes a target region of interest that matches a methylated cytosine at a CpG site within a reference nucleic acid sequence.
[0020] In one embodiment, the nucleic acid probe is a mixture of chemically or enzymatically modified methylated nucleic acid probes or unmethylated nucleic acid probes.
[0021] In one embodiment, one or more cytosines in the CG context of the probed converted nucleic acid are converted to thymine, and all cytosines in the CH context of the probed converted nucleic acid are converted to thymine.
[0022] In one embodiment, the conversion of unmethylated cytosine to uracil involves a series of TET / APOBEC enzyme conversions.
[0023] In one embodiment, the conversion of unmethylated cytosine to uracil involves TAPS.
[0024] In a second aspect, a method for determining a targeted methylation pattern within a nucleic acid molecule of a biological sample from a subject is provided, the method comprising: a) ligating a nucleic acid adapter containing a unique molecular identifier to cfDNA, wherein the cfDNA contains an unconverted nucleic acid; b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with a nucleic acid probe complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid with the reference nucleic acid sequences of the pre-identified panel of CpG or CH loci to determine the methylation profile of a cell-free DNA (cfDNA) sample derived from a subject.
[0025] In one embodiment, the step of determining the nucleic acid sequence of the converted nucleic acid includes duplex-like error correction.
[0026] In one embodiment, the nucleic acid adapter is a conversion-resistant adapter that contains each of the bases guanine, thymine, adenine, and cytosine but does not contain a 5mC-containing base or a 5hmC-containing base.
[0027] In one embodiment, the pre-identified panel of CpG or CH loci includes loci associated with transcription factor start sites.
[0028] In one embodiment, the targeted methylation pattern includes hemimethylated CpG loci.
[0029] In a third aspect, a method for determining the methylation profile of a cell-free DNA (cfDNA) sample derived from a subject is provided, the method comprising: a) ligating a nucleic acid adapter containing a unique molecular identifier to the cfDNA, wherein the cfDNA contains an unconverted nucleic acid; b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with a nucleic acid probe complementary to the pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid at a depth of more than 100x; f) determining a methylation profile of a cell-free DNA (cfDNA) sample derived from a subject, by comparing the nucleic acid sequence of the converted nucleic acid with a reference nucleic acid sequence of the pre-identified panel of CpG or CH loci
[0030] In one embodiment, the nucleic acid adapter is a conversion-resistant adapter that contains each of the bases guanine, thymine, adenine, and cytosine, but does not contain a 5mC-containing base or a 5hmC-containing base.
[0031] In one embodiment, the unique molecular identifier is 4 bp to 6 bp in length and has a 5'-thymidine overhang.
[0032] In one embodiment, the nucleic acid adapter further comprises a unique dual index (UDI) sequence. In one embodiment, the length of the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp.
[0033] In one embodiment, the method further comprises identifying a cfDNA sample derived from a tissue, identifying somatic mutations in the cfDNA sample, estimating nucleosome positions in the cfDNA sample, identifying variable methylation regions in the cfDNA sample, or identifying haplotype blocks in the cfDNA sample.
[0034] This specification further provides a method for methylation sequencing by duplex sequencing. Duplex sequencing is a tag-based error correction method that can improve sequencing accuracy, such as methylation sequencing accuracy. In this method, adapters are ligated onto nucleic acid templates and amplified using PCR. In one embodiment, the adapter includes a primer sequence and a random 12bp index. Deep sequencing provides consensus sequence information from all unique molecular tags. Based on the molecular tags and sequencing primers, a double sequence can be aligned to determine the true sequence of the DNA. Advantages of duplex sequencing include a very low error rate and the detection and removal of PCR amplification errors. In duplex sequencing, no additional library preparation steps are required after the addition of the adapter.
[0035] In one embodiment, a method for performing methylation sequencing on nucleic acid molecules of a biological sample is a) a step of preparing a methylation sequencing library from cfDNA fragments of nucleic acid molecules, comprising i) ligating a double-stranded adapter to the cfDNA fragment, ii) ligating a dual unique molecular identifier to the cfDNA fragment, and iii) preparing a methylation sequencing library from the cfDNA of the nucleic acid molecule by converting non-methylated cytosine in the cfDNA fragment to uracil using a minimally disruptive conversion method, b) a step of enriching the methylation sequencing library for sequences corresponding to CpG or CH loci, thereby generating an enriched methylation sequencing library, c) a step of sequencing the enriched methylation sequencing library to a depth of more than 100x using single-end reads or paired-end reads, thereby generating sequencing fragments of single-end reads or paired-end reads d) For each sequencing fragment of the paired-end reads, a step of correcting sequencing errors within the overlapping region of the paired-end reads; e) A step of folding the sequencing fragments into chain read families to correct errors resulting from PCR and sequencing; f) A step of folding the chain read families into double read families to identify methylation mismatches with respect to the estimated methylation status of symmetric CpG loci in the nucleic acid molecule.
[0036] In one embodiment, the minimally disruptive conversion includes enzymatic conversion, TAPS, or CAPS.
[0037] In a fourth aspect, a method for generating a classifier is provided, the method comprising: a) ligating a nucleic acid adapter containing a unique molecular identifier to each nucleic acid molecule of a biological sample from a healthy subject and a biological sample from a cancer subject; b) a step of converting unmethylated cytosine to uracil within the nucleic acid molecule using a minimally disruptive conversion method, thereby generating a converted nucleic acid; c) a step of amplifying the converted nucleic acid by polymerase chain reaction, thereby generating an amplified converted nucleic acid; d) a step of probing the amplified converted nucleic acid with a nucleic acid probe complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby generating a probed converted nucleic acid; e) a step of determining the nucleic acid sequence of the probed converted nucleic acid at a depth of more than 100x; f) a step of comparing the nucleic acid sequence of the probed converted nucleic acid with the reference nucleic acid sequences of the pre-identified panel of CpG or CH loci to obtain a set of measurements of input features representing the methylation profiles of healthy and cancer subjects; g) a step of training a machine learning model to generate a classifier for discriminating between healthy and cancer subjects.
[0038] In one embodiment, the pre-identified panel of CpG or CH loci includes loci associated with transcription start sites.
[0039] In one embodiment, the method further includes determining hemimethylated CpG or CH loci.
[0040] In one embodiment, the method further includes identifying the tissue origin of the nucleic acid molecule.
[0041] In one embodiment, the method further includes identifying the genomic location and fragment length of the nucleic acid molecule.
[0042] In one embodiment, the unique molecular identifier is 4 bp to 6 bp in length and has a 5'-thymidine overhang.
[0043] In one embodiment, the nucleic acid adapter further includes a unique dual index (UDI) sequence. In one embodiment, the length of the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp.
[0044] In one embodiment, the amplification of the converted nucleic acid includes using a primer that includes a unique dual index (UDI) sequence. In one embodiment, the length of the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp.
[0045] In one embodiment, the input features include base-wise methylation percentages for CpG, base-wise methylation percentages for CHG, base-wise methylation percentages for CHH, the number or ratio of fragments with different numbers or ratios of methylated CpGs in a region, conversion efficiency (e.g., 100 average methylation percentage for CHH), hypomethylated blocks, methylation levels of CpG, methylation levels of CHH, methylation levels of CHG, fragment length, fragment midpoint, methylation levels of chrM, methylation levels of LINE1, methylation levels of ALU, dinucleotide coverage (e.g., normalized coverage of dinucleotides), uniformity of coverage (e.g., for unique CpG sites at 1x and 10x average genomic coverage (e.g., for s4 runs)), overall average CpG coverage (e.g., depth), and average coverage in CpG islands, CGI shelves, and CGI shores.
[0046] In a fifth aspect, a classifier for discriminating between healthy individuals and cancerous individuals is provided, the classifier including a set of measurements representing methylation profiles from methylation sequencing data of healthy and cancer subjects respectively, the measurements being used to generate a set of features corresponding to characteristics of the methylation profiles, the set of features being input into a machine learning or statistical model, the machine learning or statistical model providing a feature vector useful as a classifier for discriminating between a population of healthy individuals and a population of cancerous individuals.
[0047] In a sixth aspect, a method for detecting cancer from a population of subjects is provided, the method comprising: a) assaying nucleic acids from a biological sample from a subject by using minimally invasive targeted bisulfite sequencing to obtain a methylation profile of the nucleic acids; and b) classifying the biological sample by inputting the methylation profile into a training algorithm that classifies samples of healthy and cancer subjects. c) When the training algorithm classifies a biological sample as negative for cancer with a specific confidence value, outputting a report identifying the biological sample as negative for cancer to a computer screen.
[0048] In one example, the cancer is colorectal cancer.
[0049] In a seventh aspect, the disclosure of the present invention provides a system for classifying an individual based on a methylation state, the system comprising a) A computer-readable medium product having a classifier, the classifier comprising a set of measurements representing a methylation profile from methylation sequencing data of healthy and cancer subjects, the measurements being used to generate a set of features corresponding to characteristics of the methylation profiles of healthy and cancer subjects, the set of features being input into a machine learning or statistical model, the machine learning or statistical model providing a feature vector useful as a classifier for discriminating between a population of healthy individuals and a population of cancer individuals; a computer-readable medium product; b) One or more processors for executing instructions stored on the computer-readable medium product.
[0050] In one example, the system comprises a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first-degree polynomial kernel support vector machine classifier, a second-degree polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and a non-negative matrix factorization (NMF) prediction algorithm classifier.
[0051] In one embodiment, the system comprises means for performing any of the above methods.
[0052] In one embodiment, the system comprises one or more processors configured to execute any of the above methods.
[0053] In one embodiment, the system comprises modules configured to execute respective steps of any of the above methods.
[0054] In another aspect, the disclosure of the present invention provides a method for monitoring the minimal residual disease state of a subject previously treated for a disease, the method comprising determining a methylation profile as a baseline of the methylation state as described herein, and repeating the analysis to determine the methylation profile at one or more predetermined time points, wherein a change from the baseline indicates a change in the minimal residual disease state at baseline in the subject.
[0055] In another aspect, the disclosure of the present invention provides a method for monitoring the minimal residual disease state of a subject previously treated for a disease, the method comprising a) determining a baseline of a methylation profile of a biological sample obtained from a subject at a baseline methylation state; b) determining a test methylation profile of a biological sample obtained from the subject at one or more predetermined time points after the baseline methylation state; c) determining a change in the test methylation profile when compared to the baseline of the methylation profile, the change indicating a change in the minimal residual disease state of the subject.
[0056] In some embodiments, the disease is cancer. In some embodiments, the disease is colorectal cancer.
[0057] In another aspect, the disclosure of the present invention provides a method for monitoring the minimal residual disease state of a subject previously treated for colorectal cancer, the method comprising detecting a methylated fragment in a biological sample derived from the subject, wherein the methylated fragment in the biological sample indicates a change in the minimal residual disease state at baseline for the subject's colorectal cancer.
[0058] In some embodiments, the minimal residual disease state is selected from response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.
[0059] In another aspect, a method for determining a response to treatment of a subject is provided.
[0060] In another aspect, a method for monitoring tumor burden of a subject is provided.
[0061] In another aspect, a method for detecting residual tumor after surgery of a subject is provided.
[0062] In another aspect, a method for detecting recurrence in a subject is provided.
[0063] In another aspect, a method for use as secondary screening of a subject is provided.
[0064] In another aspect, a method for use as primary screening of a subject is provided.
[0065] In another aspect, a method for monitoring cancer progression of a subject is provided.
[0066] In another aspect, the disclosure of the present invention provides a kit for tumor detection, comprising reagents for carrying out the aforementioned methods and instructions for detecting tumor signals, such as methylation signals. Examples of reagents include primer sets, PCR reaction components, sequencing reagents, minimally disruptive conversion reagents, and library preparation reagents. BRIEF DESCRIPTION OF THE DRAWINGS
[0067]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0068] Provided herein are methods that enable the improvement of library preparation and sequencing of methylated regions for methylation profiling of cfDNA. This method addresses the limitations of conventional methylation sequencing and profiling of nucleic acids in biological samples by improving coverage, coverage uniformity, resolution, and methylation data accuracy to support practical applications. The sequencing data obtained from the methods provided herein is useful for practical applications that use methylation profiling data to classify or stratify populations of individuals. Such classification and stratification of populations of individuals can include the identification of individuals with a disease, the staging of disease progression, or the response to a particular treatment for the disease.
[0069] I. Definitions As used herein, terms in the singular, such as "a," "an," and "the," include both the singular and the plural, unless the context clearly indicates otherwise.
[0070] The terms "cell-free DNA in plasma," "circulating free DNA," "cell-free DNA," or "cfDNA" may refer to DNA molecules that circulate in the cell-free portion of blood. Nucleic acids that circulate in the blood originate from necrotic or apoptotic cells, and very high amounts of nucleic acids are observed from apoptosis in diseases such as cancer. Circulating DNA in cancer is associated with prominent signs of the disease, including mutations in cancer genes and microsatellite changes. These circulating DNAs may also be referred to as circulating tumor DNA (ctDNA). Viral genomic sequences, DNA, or RNA in plasma are potential biomarkers for diseases.
[0071] In some embodiments, the cell-free fraction of blood is preferably serum or plasma. As used herein, the term "cell-free fraction" of a biological sample refers to a fraction of the biological sample that is substantially free of cells. As used herein, the term "substantially cell-free" may refer to a preparation from a biological sample containing less than about 20,000 cells / ml, less than about 2,000 cells / ml, less than about 200 cells / ml, or less than about 20 cells / ml. Genomic DNA (gDNA) refers to non-fragmented DNA released from white blood cells that contaminate fractions that are free of blood cells. To reduce gDNA-contaminated samples, a highly controlled sample processing workflow may be implemented and the sample may be screened for the presence of gDNA.
[0072] As used herein, the term "diagnosing" or "diagnosis" with respect to a condition or outcome includes predicting or diagnosing the condition or outcome, determining the predisposition to the condition or outcome, monitoring a patient's treatment, a patient's treatment response, prognosis, progression, and diagnosis of response to a particular treatment of the condition or outcome.
[0073] As used herein, the term "position" refers to the nucleotide position of a strand identified within a nucleic acid molecule.
[0074] As used herein, the term "nucleic acid" refers to DNA, RNA, DNA / RNA chimeras or hybrids, which may be single-stranded (ss) or double-stranded (ds). Nucleic acids may be genomic, derived from the genome of eukaryotic or prokaryotic cells, synthesized, cloned, amplified, or reverse transcribed. In certain embodiments of the methods and compositions, the nucleic acid preferably refers to genomic DNA as determined from the context.
[0075] As used herein, unless otherwise specified, the term "modified cytosine" refers to 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), formyl-modified cytosine, carboxy-modified cytosine, 5-carboxylcytosine (5caC), or cytosine modified with any other chemical group.
[0076] As used herein, the terms "methylcytosine dioxygenase", "dioxygenase", or "oxygenase" refer to enzymes that convert 5mC to 5hmC. Non-limiting examples of methylcytosine dioxygenases include TET1, TET2, TET3, and Negleria TET. TET2 is an example of a methylcytosine dioxygenase that oxidizes at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, or at least 99% of total 5mC.
[0077] As used herein, the terms "conversion-resistant adapter" or "conversion-resistant primer" refer to nucleic acid molecules used as adapters or primers, respectively. Instead of incorporating modified nucleotide bases to prevent base conversion, conversion-resistant adapters or conversion-resistant primers incorporate only unmodified bases, enabling total base conversion during the conversion reaction for methylation sequencing. "Unmodified bases" in the adapter / primer DNA sequence refer to conventional guanine, adenine, cytosine, and thymine.
[0078] As used herein, the term "cytidine deaminase" refers to an enzyme that deaminates cytosine (C) to form uracil (U). Non-limiting examples of cytidine deaminases include the APOBEC family of cytidine deaminases, such as APOBEC3A. In any embodiment, the cytidine deaminase described herein may have a sequence that is at least 90% identical (e.g., at least 95% identical) to the amino acid sequence of GenBank accession number AKE33285.1, which accession number is the sequence of human APOBEC3A. In some embodiments, the cytidine deaminase described herein converts unmodified cytosine to uracil with an efficiency of at least 95%, 98% or 99%, preferably at least 99%.
[0079] As used herein, the term "glucosyltransferase" or "GT" refers to an enzyme that catalyzes the transfer of a β-D-glucosyl or α-D-glucosyl residue from UDP-glucose to a 5hmC residue to form 5ghmC. APOBEC can convert 5hmC to U at a lower rate than it converts C or 5mC to U. An example of a GT is T4-beta GT (βGT). In one example, the GT may be used simultaneously with a dioxygenase. This combination reliably blocks the deamination of 5hmC, such that less than 5%, less than 3%, or less than 1% of 5hmC is converted to U by the deaminase. In another example, the GT may be used together with a dioxygenase in the same reaction mixture with DNA, such that the dioxygenase converts 5mC to 5hmC and 5caC, and the GT converts any remaining 5hmC to 5ghmC to ensure that only cytosine is deaminated.
[0080] As used herein, the terms "portion" and "aliquot" are intended to have the same meaning with respect to a nucleic acid sample and can be used interchangeably.
[0081] As used herein, the term "comparing" refers to analyzing two or more sequences relative to each other. Optionally, the comparison may be performed by aligning two or more sequences relative to each other, whereby the nucleotides located at corresponding positions are aligned relative to each other.
[0082] As used herein, the term "reference sequence" refers to the sequence of a fragment being analyzed. The reference sequence can be obtained from a public database or determined separately as part of an experiment. Optionally, the reference sequence may be hypothetical, whereby the reference sequence can be computationally deaminated (i.e., changing Cs to Us or Ts, etc.) to enable the performance of sequence comparisons.
[0083] As used herein, the terms "G", "A", "T", "U", "C", "5mC", "5fC", "c5aC", "5hmC", and "5ghmC" refer to guanidine (G), adenine (A), thymine (T), uracil (U), cytosine (C), 5-methylcytosine, 5-formylcytosine, 5-carboxylcytosine (5caC), 5-hydroxymethylcytosine, and 5-glucosylhydroxymethylcytosine, respectively. For clarity, C, 5fc, 5caC, 5mc, and 5ghmC are each distinct moieties.
[0084] The term "minimal residual disease" or "MRD" refers to a small number of cancer cells remaining in the body after cancer treatment. MRD testing can be performed to determine whether the cancer treatment has been effective and to plan further treatment. To evaluate MRD, various metrics can be used, including but not limited to response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.
[0085] The term "next-generation sequencing" or "NGS" generally applies to sequencing libraries of genomic fragments less than 1 kb in size.
[0086] As used herein, the term "healthy" refers to a subject without disease, or a sample derived therefrom. Health is a dynamic state, but the term may also refer to the pathological state of a subject lacking a reference disease state, such as cancer. In one example, when referring to a methylation profile that classifies cancer subjects, the term "healthy" refers to an individual lacking a cancer such as CRC. Other diseases or health conditions may be present in this subject, but the term "healthy" may indicate the absence of the specified disease in order to compare or classify subjects with and without the disease state and samples derived from these subjects.
[0087] As used herein, the term "threshold" generally refers to a value selected to discriminate, distinguish, or differentiate between two populations of subjects. In some embodiments, this threshold discriminates methylation states between a diseased (e.g., malignant) state and a non-diseased (e.g., healthy) state. In some embodiments, the threshold discriminates the stage of a disease (e.g., stage 1, stage 2, stage 3, or stage 4). The threshold can be set according to the target disease, for example, based on an initial analysis of a training set or computationally determined for a set of inputs having known characteristics (e.g., healthy, diseased, or disease stage). The threshold can also be set for gene regions according to predicted values of methylation at specific sites. The threshold may vary for each methylation site, and data obtained from multiple sites can be combined in a final analysis.
[0088] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the teachings of this invention, but some exemplary methods and materials are described herein.
[0089] Any citation to any publication is for the purpose of showing that it was disclosed prior to the filing date of this application and should not be construed as an admission that the claims of this patent do not predate such a publication by virtue of prior invention. Further, the provided publication date may differ from the actual publication date that can be independently verified.
[0090] As will be apparent to those skilled in the art upon reading the disclosure of this invention, each of the individual embodiments described and illustrated herein has distinct components and features that can be readily separated from, or combined with, any of the features of other various embodiments without departing from the scope or spirit of this teaching. Any method recited can be performed in the order of the recited events or in any other order that is logically possible.
[0091] All patents and publications referred to herein are hereby incorporated by reference and include all sequences disclosed therein.
[0092] II. Targeted Methylation Sequencing In the method of targeted methylation sequencing, target regions in a biological sample such as cfDNA are analyzed to determine the methylation status of the target gene sequence. In some embodiments, the target region comprises consecutive nucleotides of the target region of interest, such as at least about 16 consecutive nucleotides of the target region of interest, or hybridizes to such consecutive nucleotides under stringent conditions. In various examples, targeted sequencing can be achieved using each of the techniques of hybridization capture and amplicon sequencing.
[0093] A. Hybridization Capture The hybridization methods provided herein can be used for nucleic acid hybridization in various formats, such as solution hybridization, hybridization on solid supports (e.g., Northern hybridization, Southern hybridization, and in situ hybridization on membranes, microarrays, and cell / tissue slides). In particular, this method is suitable for solution-phase hybrid capture for target enrichment of certain genomic DNA sequences (e.g., exons) used in targeted next-generation sequencing. In the hybrid capture approach, cell-free nucleic acid samples are subjected to library preparation. As used herein, "library preparation" includes end repair, A-tailing, adapter ligation, or other preparations performed on cell-free DNA to enable subsequent DNA sequencing. In one example, the prepared cell-free nucleic acid library sequences include adapters, sequence tags, and index barcodes that are ligated to cell-free nucleic acid sample molecules. To facilitate library preparation for next-generation sequencing techniques, various commercially available kits are also available. Next-generation sequencing library construction may include the step of preparing nucleic acid targets using a series of enzyme reactions adjusted to generate an aggregation of DNA fragments of a specific size for high-throughput sequencing. With the progress and development of various library preparation techniques, the application of next-generation sequencing in fields such as transcriptomics and epigenetics has been expanding.
[0094] As a result of the improvement of sequencing techniques, changes and improvements have been observed in library preparation. Examples of kits for next-generation sequencing library preparation used herein include those developed by companies such as Agilent, Bioo Scientific, Kapa Biosystems, New England Biolabs, Illumina, Life Technologies, Pacific Biosciences, and Roche.
[0095] In various examples for the targeted capture gene panel, various library preparation kits can be selected from Nextera Flex (Illumina), Ion Ampliseq (Thermo Fisher Scientific), Genexus (Thermo Fisher Scientific), Agilent ClearSeq (Illumina), Agilent SureSelect Capture (Illumina), Archer FusionPlex (Illumina), BiooScientific NEXTflex (Illumina), IDT xGen (Illumina), Illumina TruSight (Illumina), Nimblegene SeqCap (Illumina), and Qiagen GeneRead (Illumina).
[0096] In some embodiments, the hybrid capture method is performed on the prepared library sequences using specific probes. As used herein, the term "specific probe" may refer to a probe specific to a known methylation site. In some embodiments, the specific probe is designed based on using the human genome as a reference sequence and using specific genomic regions known to have methylation sites as target sequences. Specifically, the genomic regions known to have methylation sites may include at least one of the following promoter regions, CpG island regions, CGI shore regions, and imprinting gene regions. Thus, when performing hybrid capture using the specific probe of some embodiments, sequences in the sample genome complementary to the target sequence, for example, regions in the sample genome known to have methylation sites (also referred to herein as "specific genomic regions") can be efficiently captured.
[0097] According to one example, the methylation regions described herein are used to design specific probes. In some embodiments, the specific probes are designed using commercially available methods such as, for example, the eArray system. The length of the probe can be long enough to hybridize with sufficient specificity to the desired methylation region. In various examples, the probe is a decamer, undecamer, dodecamer, tridecamer, tetradecamer, pentadecamer, hexadecamer, heptadecamer, octadecamer, nonadecamer, or eicosamer.
[0098] The target regions for methylation analysis can be screened by leveraging database resources (such as gene ontology). According to the principle of complementary base pairs, a single-stranded capture probe may be combined complementarily with a single-stranded target sequence to capture the target region well. In some embodiments, the designed probe may be designed as a solid capture chip (where the probe is immobilized on a solid support) or as a liquid capture chip (where the probe is free in a liquid). However, due to limiting factors such as probe length, probe density, and high cost, solid capture chips are rarely used, while liquid capture chips are used more frequently.
[0099] In some embodiments, compared to a normal sequence (where the average content rates of bases A, T, C, and G are each 25%), a GC-rich sequence (where the average content rate of GC bases is higher than 60%) may result in a decrease in capture efficiency due to the molecular structures of bases C and G. In important research regions, such as CGI regions (CpG islands), it may be necessary to increase the amount of probe used to obtain sufficient and accurate CGI data.
[0100] B. Amplicon-based Sequencing The fragment of the converted DNA may be amplified. In some embodiments, the amplification is carried out using primers designed to anneal to a methylated conversion target sequence having at least one methylation site. By methylation sequencing conversion, unmethylated cytosine is converted to uracil, and 5-methylcytosine is not affected. A "conversion target sequence" may refer to a sequence in which cytosine known to be a methylation site is fixed as "C" (cytosine), while cytosine known to be unmethylated (not methylated) is fixed as "U" (uracil; may be treated as "T" (thymine) for the purpose of primer design).
[0101] In various examples, the source of DNA is cell-free DNA obtained from whole blood, plasma, serum, or genomic DNA extracted from cells or tissues. In some embodiments, the size of the amplified fragment is about 100 to 200 base pairs in length. In some embodiments, the DNA source is extracted from a cell source (e.g., tissue, biopsy, cell line), and the size of the amplified fragment is about 100 to 350 base pairs in length. In some embodiments, the amplified fragment contains at least one 20-base pair sequence containing at least one, at least two, at least three, or more than three CpG dinucleotides. The amplification may be carried out using a set of primer oligonucleotides according to the present disclosure, or a thermostable polymerase may be used. The amplification of multiple DNA segments may be carried out simultaneously in one and the same reaction vessel. In some embodiments, two or more fragments are amplified simultaneously. For example, the amplification may be carried out using the polymerase chain reaction (PCR).
[0102] Primers designed to target such sequences may exhibit some bias towards the converted methylation sequences. In some embodiments, the PCR primers are designed to be methylation-specific for targeted methylation sequencing applications. Methylation-specific primers may enable higher sensitivity in some applications. For example, the primers may be designed to include characteristic nucleotides (specific for the methylated sequences after bisulfite conversion) arranged to achieve optimal discrimination, for example, in PCR applications. The characteristic nucleotides may be arranged at the 3' final position or the second-to-last position.
[0103] In some embodiments, the primers are designed to amplify DNA fragments with a length of 75 - 350 bp, which is the general size range of circulating DNA. By optimizing primer design considering the target size, the sensitivity of the methods described herein can be enhanced. The primers may be designed to amplify regions that are about 50 - 200, about 75 - 150, or about 100 or 125 bp.
[0104] In one embodiment, the amplification step includes using primers that contain a unique dual index (UDI) sequence.
[0105] In one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0106] In some embodiments, the methylation status of preselected CpG positions within a nucleic acid sequence is detected by an amplicon-based approach using methylation-specific PCR (MSP) primer oligonucleotides. By using methylation-specific primers for the amplification of bisulfite-treated DNA, methylated nucleic acids and unmethylated nucleic acids can be distinguished. The MSP primer pair includes at least one primer that hybridizes to the converted CpG dinucleotide. Thus, the sequence of said primer contains at least one CpG, TpG, or CpA dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the 3' position of the C position in CpG. Thus, the base sequences of these primers may include sequences having a length of at least 18 nucleotides that hybridize to the pre-treated nucleic acid sequence and its complementary sequence, and the base sequences have at least one CpG, TpG, or CpA dinucleotide. In some embodiments of the method, the MSP primer has 2 to 5 CpG, TpG, or CpA dinucleotides. In some embodiments, the dinucleotide is located within the 3' half of the primer. For example, in a primer having a length of 18 bases, the specified dinucleotide is located within the first 9 bases from the 3' end of the sequence. In addition to the CpG, TpG, or CpA dinucleotide, the primer may further include a plurality of methylated bases (e.g., those in which cytosine is converted to thymine, or those in which guanine is converted to adenine on the strand to which it hybridizes). In some embodiments, the primer is designed to have two or fewer cytosine and / or guanine bases.
[0107] In some embodiments, each of the regions is amplified in multiple portions using multiple primer pairs. In some embodiments, these portions do not overlap. The portions may be adjacent or spaced apart (e.g., spaced 10, 20, 30, 40, or 50 bp apart). Since target regions (including CpG islands, CpG shores, and / or CpG shelves) are typically longer than 75 - 150 bp, in this example, the methylation status of sites across more (or all) of a given target region can be evaluated.
[0108] Primers can be designed for the target region using suitable tools such as Primer3, Primer3Plus, Primer - BLAST. As discussed, bisulfite conversion converts cytosine to uracil and 5'-methyl - cytosine to thymine. Thus, primer positioning or targeting can utilize bisulfite - converted methylation sequences depending on the degree of methylation specificity required.
[0109] III. Library Preparation for Enzymatic Methylation Sequencing In a first aspect, a method for the preparation of a sequencing library is provided. The methods described herein provide a library that is acceptable for both next - generation unmethylated and methylated sequencing applications, thereby providing sequencing data for two applications from a single sample. The resulting raw sequencing data can be used not only for analysis of methylation status but also for conventional cfDNA analysis such as copy number variation, detection of germline variants, detection of somatic variants, nucleosome positioning, transcription factor profiling, chromatin immunoprecipitation, etc.
[0110] A. Adapter Ligation for Targeted Sequencing Applications In one aspect, the method of the present invention preserves the integrity and information of nucleic acid sequences for methylation profiling. In one example, by combining dsDNA adapter ligation prior to enzymatic conversion, endpoint information of the fragments is preserved while providing the highest possible library complexity for target enrichment (or directly for genome-wide sequencing), thereby improving the sensitivity to detect rare events such as methylated ctDNA. A comparison of the advantages of adapter ligation before conversion is shown in FIG. 1.
[0111] In one example, nucleic acid adapters are ligated to the 5’ and 3’ ends of a population of nucleic acid fragments in a biological sample to generate a sequencing library. In one example, a collection of nucleic acid adapters is ligated to nucleic acid fragments in a sample, where the collection of adapters comprises aliquots of 4bp, 5bp, and 6bp unique molecular identifier (UMI) sequences to enable T / A overhang ligation, and then an invariant thymidine (T) at the last position (i.e., the 3’ end). Thus, the UMI is positioned adjacent to the inserted nucleic acid of the library. During sequencing, the UMI is also sequenced as part of the read at the 5’ end (alternatively, the UMI matches the library insert at the sequencing read level). The invariant T is staggered over three positions to maintain base diversity at the sequencing position. In contrast, using a single-length UMI with an invariant thymidine results in low-complexity sequencing at the position corresponding to the invariant thymidine, resulting in a decrease in sequencing quality. The first 4bp of each UMI comprises a set of 4bp core UMI sequences with an edit distance of 2 or more and a balanced nucleotide and color. Despite the variable-length UMI sequences, using a single-length core UMI facilitates the use of bioinformatics tools constructed for single-length UMIs for UMI extraction and duplicate removal. Thus, the 4bp core sequence functions as a recognition sequence to inform bioinformatics tools to trim 5, 6, or 7 bases (including the invariant T), thereby maintaining accurate cfDNA endpoint information. A schematic diagram showing the staggered adapters is shown in Figure 2. By using UMIs, read duplicate removal, single-strand error correction, and duplex reconstruction after sequencing are possible, thereby enabling the use of the reverse complement of the read to enhance error correction, also known as duplex error correction. In another example, a unique dual index (UDI) is an additional sequence that can be added to the UMI-containing adapter during library purification to effect sample barcoding and post-sequencing sample demultiplexing ().In various examples, the UDI array is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, or 12 bp in length.
[0112] In various embodiments, the nucleic acid adapter may include a UMI that is 4 bp to 6 bp in length and has a 5’ thymidine overhang. The UMI is designed to be non-unique (i.e., drawn from a specific set of constrained sequences).
[0113] In one embodiment, some UMIs include one or more methylcytosine bases. The efficiency of enzymatic methylation conversion reactions (including TET oxidation and APOBEC deamination) can be evaluated based on the UMI mismatch rate, which is the proportion of UMIs that do not match a specific set of designed UMI sequences. The UMI mismatch rate can be used as a quality control metric embedded to assess the quality of a sequencing library. Additionally, if a perfect UMI match is required in a bioinformatics pipeline, the UMI mismatch rate may be used as a filter to remove individual reads that may have low quality due to incomplete conversion.
[0114] In various embodiments, the UMI mismatch rate is less than 6%, less than 5%, less than 4%, less than 3%, or less than 2%.
[0115] In another embodiment, the UMI contains one or more cytosines with modifications that can be used to monitor enzymatic activity. Non-limiting examples of these modified bases include 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0116] B. Enzymatic conversion for DNA methylation sequencing applications Tet-assisted pyridine borane sequencing (TAPS) is a minimally destructive conversion methylation sequencing method for converting cytosine to uracil in nucleic acids. Since this method does not use bisulfite, DNA degradation can be minimized, and the length of nucleic acid molecules can be maintained while achieving a conversion rate equivalent to that of sodium bisulfite sequencing. TAPS can provide higher sequencing quality scores for cytosine and guanine base pairs and more uniform coverage of various genomic features such as CpG islands.
[0117] In TAPS, the ten-eleven translocation (Tet1) enzyme oxidizes both 5mC and 5hmC to 5caC. Pyridine borane reduces 5caC to dihydrouracil, which is a uracil derivative that is converted to thymine after PCR. TAPS can also be performed in two other ways: TAPSβ and chemically assisted pyridine borane sequencing (CAPS). In TAPSβ, 5hmC can be specifically detected by protecting it from oxidation-reduction reactions by labeling 5hmC with glucose using β-glucosyltransferase. In CAPS, potassium perruthenate acts as a chemical substitute for Tet1 and specifically oxidizes 5hmC, enabling direct detection.
[0118] In one example, the combination of enzymatically converting unmodified C to U and staggering UMI adapters relative to the library insert is useful for targeted sequencing of methylated libraries. In low-depth sequencing applications, since the sample cfDNA is not degraded to the same extent, this combination may enable a reduction in the quantitative input of plasma or the qualitative input of cfDNA compared to bisulfite conversion sequencing.
[0119] In high-depth sequencing applications, since the cfDNA is not degraded to the same extent, higher-depth sequencing may be obtained compared to bisulfite conversion sequencing from a similar input of plasma or cfDNA.
[0120] In one example, cytosine present in the adapter nucleic acid is modified with a 5-methyl group or a 5-hydroxymethyl group to prevent C-T conversion in the adapter.
[0121] One advantage of this approach is that adapter ligation prior to conversion maintains the endpoint and length information of the fragment compared to bisulfite conversion and subsequent ssDNA adapter ligation approaches. Considerable degradation of the nucleic acid prior to adapter ligation can result in loss of endpoint and length information of informative fragments.
[0122] Enzymatic conversion of unmodified C to U can be less burdensome on the nucleic acid fragments of the sample and can result in more complete and uniform coverage compared to the bisulfite method. Bisulfite degradation of DNA is not uniform, and some sequences are preferentially degraded over other sequences such as CG dinucleotides at precisely the sites that are being investigated for methylation sequencing. Thus, the enzymatic approach results in higher CpG site coverage than the bisulfite conversion method using the same number of unique reads and results in high uniformity of the captured reads in target enrichment applications. Additionally, non-bisulfite methods (e.g., enzymatic and chemical conversion methods such as TAPS) enhance the resolution of biological signals and specifically provide the ability to distinguish between 5mC and 5hmC methylation in nucleic acid sequences. This information and additional resolution can be beneficial in computational approaches and other methods.
[0123] In some examples, subjecting DNA or barcoded DNA to an enzymatic reaction that converts cytosine nucleobases of the DNA or barcoded DNA to uracil nucleobases includes "performing an enzymatic conversion."
[0124] In various examples, glucosylation and oxidation reactions overcome the observed specific deamination of 5hmC and 5mC by deaminase. Deaminase converts 5mC and unmodified C to U, but not 5ghmC and 5caC. Non-limiting examples of deaminase include APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like). Embodiments described herein utilize enzymes that are substantially sequence unbiased in cytosine glucosylation, oxidation, and deamination. Further, these embodiments do not substantially cause non-specific damage to DNA during glucosylation, oxidation, and deamination reactions.
[0125] In some embodiments, a glucosyltransferase (GT), such as β-glucosyltransferase (βGT), is utilized to covalently attach glucose to 5hmC and protect this modified base from deamination. Other enzymatic or chemical reactions for modifying 5hmC may be used to achieve the same effect.
[0126] Generally, and in one aspect, the methods provided herein include: (a) treating an aliquot (portion) of a nucleic acid sample with dioxygenases, such as TET2, and βGT in a reaction mixture to produce a reaction product in which substantially all modified cytosines (C) are oxidized or, in the case of 5hmC, glucosylated; and (b) treating this reaction product with a cytidine deaminase to convert substantially all unmodified Cs to Us. The term "modified" cytosine as used throughout these examples and embodiments refers to one or more of 5mC, 5hmC, 5ghmC, 5fC, and 5caC, where 5caC is produced by oxidation of 5mC, 5hmC, and 5fC to completion. βGT reacts only with 5hmC. However, some of the 5hmC is converted to 5fC by dioxygenase and further to 5caC before glucosylation occurs. In the presence of dioxygenase, most of the 5mC is oxidized to completion of 5caC, but some remaining 5hmC may be produced. However, the remaining 5hmC can be glucosylated by βGT to prevent the low deamination rate of 5hmC that may otherwise reduce the accuracy of methylation sequencing.
[0127] Thus, the described method greatly differentiates unmodified cytosine from modified cytosine by treating the nucleic acid with dioxygenase prior to deamination. However, the amount of naturally occurring 5mC in genomic DNA can substantially exceed the amount of 5hmC and, as a result, can exceed the amounts of naturally occurring 5fC and 5caC. Thus, the amount of naturally occurring modified cytosine is generally considered to be an approximation of the amount of naturally occurring 5mC.
[0128] In one embodiment, the method is adaptable for performing 5hmC sequencing. The 5hmC sequencing method further includes treating an aliquot of a nucleic acid sample with βGT in the absence of dioxygenase and then treating with cytidine deaminase to produce a reaction product in which substantially all 5hmC in the aliquot is glucosylated and substantially all unmodified C and 5mC are converted to U. After PCR amplification, U is converted to T, and thus cytosine and 5mC become indistinguishable during sequencing. By sequencing the resulting reaction product and comparing it to a reference sequence, 5hmC can be distinguished from C and 5mC. By distinguishing these sites, these modified nucleotides can be mapped to a reference sequence, such as a reference sequence from a database or a reference sequence determined independently.
[0129] In some embodiments, the reaction product or its amplification product of βGT and deaminase with dioxygenase can be sequenced to determine which C is methylated (which may contain a small fraction of 5hmC) and which C is unmodified. In some embodiments, the reaction product or its amplification product of βGT and deaminase without dioxygenase can be sequenced to determine which C is hydroxymethylated and which C is not hydroxymethylated. In some embodiments, the reaction product or its amplification product of βGT and deaminase without dioxygenase can be sequenced to determine which C is hydroxymethylated and which C is unmodified. The reference DNA can be generated by sequencing the reaction product resulting from not reacting the nucleic acid sample with any one of dioxygenase, βGT, and deaminase. Alternatively, the reference sequence is, for example, a known reference sequence from a database of sequences.
[0130] In one embodiment, the sequence of the reaction product of dioxygenase and deaminase having βGT can be compared with a reference sequence. Optionally, this can be compared with the sequence of the reaction product of βGT (without dioxygenase) and deaminase to determine which cytosines in the nucleic acid sample are modified by methyl groups versus hydroxymethyl groups.
[0131] In one aspect, a method for performing targeted methylation sequencing of a nucleic acid sample is provided, the method comprising: a) ligating a nucleic acid adapter containing a unique molecular identifier to cfDNA, wherein the cfDNA contains unconverted nucleic acid; b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with a nucleic acid probe complementary to a panel of pre-identified CpG or CH loci to enrich for sequences corresponding to the panel of pre-identified CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid at a depth of more than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid with a reference nucleic acid sequence of a panel of pre-identified CpG or CH loci to determine the methylation profile of a cell-free DNA (cfDNA) sample from a subject.
[0132] If the test converted nucleic acid sequence is a T corresponding to the reference C at a specified CpG locus, the C was not methylated in the original test nucleic acid fragment. In contrast, if both the test converted nucleic acid sequence and the reference sequence are C at a specified CpG locus, the C was methylated in the original test nucleic acid fragment.
[0133] In one embodiment, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 50 - 500 fold, about 25 - 1000 fold, about 50 - 500 fold, about 250 - 750 fold, about 500 - 200 fold, about 750 - 1500 fold, or about 100 - 2000 fold. In some embodiments, the nucleic acid sequence is sequenced at a depth of 100 fold or greater than 500x.
[0134] In one example, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 500 fold, about 1000 fold, about 2000 fold, about 3000 fold, about 4000 fold, about 5000 fold, about 6000 fold, about 7000 fold, about 8000 fold, about 9000 fold, about 10000 fold, or greater than 5000 fold.
[0135] In one example, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 300 fold unique, about 400 fold unique, about 500 fold unique, about 600 fold unique, about 700 fold unique, about 800 fold unique, about 900 fold unique, or about 1000 fold unique, or greater than 500 fold unique.
[0136] C. Target Enrichment Sequencing Applications Furthermore, in a target capture application during sequencing, a method for enriching a desired methylation region is provided. A potential problem when applying a target enrichment capture panel using a DNA methylation library is the low rate of on-target reads / high rate of off-target DNA fragment capture. For each region within the panel, probes can be designed to target DNA derived from methylated CpGs or DNA derived from unmethylated CpGs. In either probe type, all CpG sites along the region are considered to be either unmethylated or methylated as required for the probe type. The probes can hybridize to library molecules after bisulfite / enzyme conversion and PCR amplification. Thereafter, only the library molecules captured by the probes are sequenced. In this method, only a very small portion of the genome is sequenced, thus having the advantage of reducing the sequencing cost. In one example, about 0.1% of the genome is sequenced. In one example, about 0.3% of the genome is sequenced. In one example, about 0.5% of the genome is sequenced. In one example, about 0.7% of the genome is sequenced. In one example, about 1% of the genome is sequenced. In other examples, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the genome is sequenced.
[0137] In both bisulfite and enzyme conversion libraries, a target capture enrichment approach can result in significant off-target capture rates. The off-target capture rate is partially due to all cytosines not at CpG sites being converted from C to T in both types of probes that hybridize to DNA derived from methylated CpGs. The decrease in cytosine content in the probes leads to a decrease in sequence complexity and thus a decrease in the specificity of the probes that hybridize to target library molecules.
[0138] As used herein, the terms "methylated probe" and "unmethylated probe" refer to probes used to hybridize to methylated CpGs and unmethylated CpGs, respectively, in a converted nucleic acid sequence. The probe may be designed to recognize the converted nucleic acid sequence. In a converted methylated CpG probe, C remains C after conversion. In a converted unmethylated CpG probe, C is converted to T after conversion. In both the converted methylated probe and the unmethylated probe, all Cs of non-CpG dinucleotides are converted to T after conversion.
[0139] The methylated probe retains some cytosines (i.e., the cytosines of the CpG sites). In contrast, in the unmethylated probe, all cytosines are converted to thymines. The unmethylated probe is less complex than the methylated probe and is likely to contribute preferentially to the off-target capture rate. In one example, a probe that hybridizes to DNA derived from methylated CpGs is used in a target enrichment method. In one example, a probe having a sequence substantially complementary to a target that hybridizes to DNA derived from methylated CpGs is used in a target enrichment method.
[0140] Probes that hybridize to DNA derived from methylated CpGs for target enrichment can be selected to achieve different embodiments. The target capture hybridization reaction occurs at a single temperature. However, the optimal melting temperature (Tm) of a probe that hybridizes to DNA derived from methylated CpGs is, on average, higher than the Tm of a probe not designed to hybridize to DNA derived from methylated CpGs.
[0141] The base pair of cytosine contains three hydrogen bonds, while the base pair of thymine contains two hydrogen bonds. Converting cytosine in the probe to thymine reduces the hydrogen bonds, thus lowering the Tm of the probe. Since the methylated probe contains some cytosine and the unmethylated probe does not retain cytosine, the Tm of the methylated probe is higher than that of the corresponding unmethylated probe. When the number of CpG sites increases in a region, the difference in melting temperature between the methylated probe and the unmethylated probe also increases. Probes with a high melting temperature may hybridize to the target DNA fragment more efficiently than probes with a low melting temperature. Generally, the hybridization temperature is selected to be relatively high to promote on-target capture. However, at the general hybridization temperature, the methylated probe has a high melting temperature due to the retention of multiple cytosines, so it hybridizes more efficiently than the unmethylated probe. A high melting temperature may introduce a bias such that the % of CpG methylation level measured by the target capture hybridization approach is higher compared to the level measured by sequencing of the pre-capture library.
[0142] In one example, only a single type of methylated or unmethylated probe is used in the hybridization reaction to enrich for hypermethylated or hypomethylated library molecules, respectively. By using a single type of methylated or unmethylated probe, the problem of different melting temperatures between probe types can be avoided. Using a single type of probe can also more efficiently capture (or enrich) the same DNA fragment type. In one example, using only methylated probes results in preferential binding of hypermethylated ROIs over hypomethylated ROIs. In another example, enrichment of unmethylated ROIs is obtained by using only unmethylated probes.
[0143] By using only a single probe type, higher hybridization temperatures can also be used to reduce off-target capture without affecting the relative balance of capturing methylated and unmethylated ROIs. Thus, the probe panel can be designed based on the desire to enrich hypermethylated or hypomethylated DNA fragments. As an example, when quantification of both hypermethylated and hypomethylated DNA fragments is desired, two parallel but separate hybridization reactions are employed for both methylation states.
[0144] D. Methylation Analysis In various examples, once enzymatic methylation sequencing is complete, an assay is used to analyze the methylation state of nucleic acids in a biological sample. In one example, whole-genome enzymatic methyl-sequencing (“WG EM-seq”) provides high-resolution sequencing by characterizing DNA methylation of almost all cytosine nucleotides within the genome. Other targeted methods, such as targeted enzymatic methyl-sequencing (“TEM-seq”), may be useful for methylation analysis.
[0145] In other examples, assays that have conventionally been used for bisulfite conversion can be employed with minimally disruptive conversion methods such as enzymatic conversion, TAPS, and CAPS. In various examples, assays used for methylation analysis can be mass spectrometry, methylation-specific PCR (MSP), reduced representation bisulfite sequencing (RRBS), HELP assay, GLAD-PCR assay, ChIP-on-chip assay, restriction landmark genome scan, methylated DNA immunoprecipitation (MeDIP), pyrosequencing of bisulfite-treated DNA, molecular break light assay, methylation-sensitive Southern blotting, high-resolution melting (HRM or HRMA), ancient DNA methylation reconstruction, or methylation-sensitive single nucleotide primer extension assay (msSNuPE).
[0146] The methylation profile of cfDNA can be identified by applying a sequence alignment method to map methyl-sequencing reads obtained from whole-genome or targeted methyl-sequencing of the human reference genome. Non-limiting examples of sequence alignment methods include bwa-meth, bismark, Last, GSNAP, BSMAP, NovoAlign, Bison, metagenomic phylogenetic analysis (e.g., MetaPhlAn2), BLAT, Burrows-Wheeler Aligner (BWA), Bowtie, Bowtie2, Bfast, BioScope, CLC bio, Cloudburst, Eland / Eland2, GenomeMapper, GnuMap, Karma, MAQ, MOM, Mosaik, MrFAST / MrsFAST, PASS, PerM, RazerS, RMAP, SSAHA2, Segemehl, SeqMap, SHRiMP, Slider / SliderII, Srprism, Stampy, vmatch, ZOOM, and SOAP / SOAP2 alignment tools.
[0147] CpG error correction using E. duplex UMI-based methylation consensus calls Methylation analysis involves analyzing sequencing data based on whether the "C" within the CpG context is read as "C" (methylated) or "T" (unmethylated) in sequencing. However, "T" may appear at these positions in sequencing for reasons other than the presence of unmethylated CpGs in the parental DNA molecule. These reasons include sequencing errors, PCR errors, nucleotide fill-in during end repair, DNA damage, germline single nucleotide polymorphisms (SNPs) that replace CpG with another dinucleotide, somatic mutations that replace CpG with another dinucleotide, and overconversion (i.e., 5mC is converted to T despite the presence of a methylation mark). In addition, "C" may appear at these positions during sequencing for reasons other than the presence of methylated CpGs in the parental DNA molecule. These reasons include sequencing errors, PCR errors, DNA damage, and incomplete conversion (unmethylated C is not converted to T despite the absence of a methylation mark), etc. If most of these error modes cannot be corrected, the methylation status of CpGs cannot be accurately read, and the detection of rare events (e.g., ctDNA molecules in early cancer) that require extremely accurate reading of the methylation status of CpGs may be limited. In addition, methods that cannot consider the information of the duplex structure cannot distinguish between hemimethylated CpG sites and symmetrically methylated CpG sites. Such information is useful for interpreting the biological significance of methylation signals. For example, hemimethylation can directly identify de novo methylation events, thereby distinguishing between de novo factors and maintenance factors.
[0148] The duplex sequencing approach overcomes the limits of sequencing accuracy by addressing these extensive errors. For example, duplex sequencing reduces errors by independently tagging and sequencing each of the two strands of a DNA duplex. Since the two strands are complementary, true mutations can be found at the same position on both strands. Similarly, since CpG dinucleotides are symmetric, fully methylated CpG motifs have methylated cytosines at opposing adjacent positions on both strands. In contrast, PCR or sequencing errors occur on only one strand. This method can overcome the technical limitations of methods that utilize data from single strands because it uniquely utilizes the redundant and additional information that exists between the strands of double-stranded DNA.
[0149] For enzymatic methylation sequencing, the efficiency of APOBEC conversion of individual fragments can be evaluated by the number of cytosines in the CHH context that are sequenced as cytosine. In the case of an APOBEC reaction with 100% efficiency, all cytosines in the CHH context are converted to uracil and sequenced as thymine. In contrast, cfDNA fragments in which the APOBEC enzyme did not act efficiently (i.e., the conversion was incomplete) may contain one or more cytosines that were not converted to uracil within the CHH context, which are sequenced as cytosine. The number of unconverted cytosines in the CHH context can be used as a filter to remove reads that may be unreliable and noisy due to incomplete conversion.
[0150] Many enzymes that act on nucleic acids have sequence preferences that introduce biases in terms of which sites are efficiently acted upon by the enzyme. Using experimental data, the sequence preferences of individual enzymes can be identified. In various embodiments, this data can be used to mask potential sites that are likely to be incompletely converted by the enzyme. As an example, APOBEC A3A has 12-fold discrimination for cytosines preceded by A compared to cytosines preceded by T.
[0151] In one example, a method for duplex methylation consensus calling is provided, the method comprising: a) preparing a methylation sequencing library from cfDNA using enzymatic conversion, comprising: (i) ligating double-stranded adapters to nucleic acid fragments obtained from a biological sample; (ii) performing target enrichment to enable ultra-deep sequencing of a desired locus pre-identified; (iii) preparing a methylation sequencing library from cfDNA using enzymatic conversion (such that neither strand is damaged), and ligating duplex UMI prior to enzymatic conversion for tagging the duplex prior to the denaturation step involved in enzymatic conversion; and b) a target enrichment step to enable ultra-deep sequencing of a specific desired locus; c) sequencing the enriched library using single-end reads or paired-end reads; d) correcting sequencing errors within the overlapping region of paired-end reads for the sequenced fragments of paired-end reads; e) folding the sequenced fragments into chain read families to correct errors arising from PCR and sequencing; f) folding the chain read families into double read families to identify inconsistent methylation states of symmetric CpGs.
[0152] A schematic diagram of this method is shown in Figure 3.
[0153] CpG "error correction" of methylated array data using duplex information provides the advantage of filtering out noise that could otherwise reduce the sensitivity or specificity of classifiers that use methylated array data as input. Due to nucleotide imbalance being introduced into the array after conversion, a unique UMI design that uses methylated UMIs placed alternately in the context of methylation sequencing enhances sequencing accuracy and can increase data output by assisting in the identification of clusters (using platforms such as the NextSeq sequencer in particular), reducing the dependence on adding large amounts of PhiX data (increasing the depth of sequencing and reducing associated costs). Different from standard duplex sequencing that analyzes the concordance of variant calls in base-paired nucleotides, CpG duplex-based duplex sequencing methods evaluate the symmetry of inter-strand CpG methylation (1bp offset). In certain examples, duplex sequencing can also distinguish SNPs from unmethylated CpGs. Enzymatic methylation sequencing methods have the advantage over bisulfite-based methods in that they can efficiently capture the sequences of both strands.
[0154] F. CpG error correction using a conversion-resistant adapter for duplex UMI-based methylation consensus calls in enzymatic methylation sequencing In another aspect, a conversion resistance adapter and a primer are used for methylation sequencing. Sequencing methods used to identify the position of base modifications, such as bisulfite sequencing or enzymatic methylation sequencing (EM-seq) used to identify 5mC, operate by chemically or enzymatically changing each unmodified cytosine base (C) to change the base pairing properties of C. For example, during the EM-seq process, all unmodified Cs are converted to uracil (U) by the APOBEC enzyme and then sequenced as thymine (T). The bases of 5mC are not converted and are sequenced as C. Since the bases can only be converted when the DNA is single-stranded, the double-stranded DNA must be denatured before the conversion reaction from C to U.
[0155] One of the potential problems that can occur when duplex sequencing and methylation sequencing are combined is the reduction in PCR amplification and sequencing. Since the adapter must be ligated onto the DNA while the DNA is still double-stranded (i.e., before base conversion), all Cs in the adapter are converted to U, thereby preventing efficient PCR amplification and sequencing.
[0156] A solution to this problem is to use an adapter with a modified base that is not converted or is less likely to be converted during the deamination reaction (e.g., 5mC, 5hmC, or other C variants). However, oligonucleotides containing modified bases are often significantly more expensive than oligonucleotides containing only standard bases. Furthermore, this solution is generally only effective for bisulfite methylation sequencing where 5mC cannot be converted.
[0157] Unlike bisulfite sequencing, the EM-seq process requires an additional enzymatic step necessary to prevent the conversion of 5mC or 5hmC to U by APOBEC. This step uses either a Tet enzyme that oxidizes the base of 5mC or 5hmC, or βGT that glucosylates 5hmC, thereby protecting 5hmC from conversion. If this step is not completely efficient, some of the 5mC or 5hmC in the adapter will be converted to uracil, resulting in loss of library complexity and degradation of sequencing quality. The Tet oxidation reaction is very sensitive to reaction conditions and can vary the quality of the sequencing library.
[0158] To improve the robustness of EM-seq duplex sequencing against oxidation efficiency (and to reduce the current economic burden), conversion-resistant adapters containing only unmodified bases can be used. Unmodified bases refer to the conventional bases guanine, cytosine, adenine, and thymine in an unmodified state. Contrary to the conventional method of restricting the total conversion of adapter molecules using modified bases such as 5mC and 5hmC, this approach allows all cytosine conversions in the adapter to achieve improved efficiency and sequencing quality. An example of a conversion-resistant adapter is shown in Panel A of Figure 4.
[0159] Sequencing libraries generated using these conversion-resistant adapters can be amplified and sequenced using a set of PCR and sequencing primers that match the original adapter sequence. After conversion, the sequencing library can be amplified and sequenced using PCR and sequencing primers that match the converted adapter sequence in Panel B of Figure 4.
[0160] G. Use of Internal Process Control during Enzymatic Conversion For targeted enzymatic methylation sequencing, synthetic internal process control (IPC) can be used to monitor the oxidation and deamination reactions during enzymatic methylation conversion.
[0161] In various embodiments, the IPC may include all 256 possible cytosine contexts in a window of two bases before and after C (NNCNN).
[0162] In various embodiments, the IPC is a duplex synthesized by PCR that includes either 100% unmodified C, 100% methylated C, or 100% hydroxylated C (or another modification to C). In this regard, the conversion or protection efficiency of the IPC can be monitored. In some embodiments, the conversion or protection efficiency can be monitored by sequencing or quantitative PCR.
[0163] H. Hemimethylation analysis In another example, the use of UMIs in the methylated sequence enables error correction and analysis / removal of hemimethylation. Alternatively, strand-specific methylation sequencing allows for the identification of hemimethylated DNA. The methylation status of CpG / CpG dyads is usually consistent, i.e., either fully methylated or fully unmethylated. However, CpG / CpG dyads with inconsistent methylation status, i.e., hemimethylated, generally occur at low to moderate frequencies, except in regions where transcriptional silencing or reactivation is occurring or temporarily during DNA replication. Such hemimethylated dyads can provide additional information that can inform classifiers when stratifying populations. Recognizing hemimethylated dyads allows for a more complete methyl sequence profile to be obtained, and the option to remove or include this information when generating classifiers.
[0164] Another advantage of the enzymatic methylation sequencing approach is that it can better distinguish methylated C from unconverted C. By maintaining the integrity and length of the fragments through enzymatic conversion, the accuracy of determining the true methylation state of nucleic acid molecules can be enhanced using duplex UMI methyl sequences. This method can account for errors that may occur, for example, during extraction (DNA damage), library preparation (end repair fill-in), enzymatic conversion (under-conversion or over-conversion), PCR (base incorporation errors), and sequencing (base call errors). Improving the accuracy of methylation state determination improves the characterization and generation of classifiers for stratifying populations using these methylation-based epigenetic sequence differences. In one example, adapter directionality is utilized to distinguish dsDNA fragments derived from the top strand versus the bottom strand (based on which genomic strand read1 maps to), which is schematically shown in Figure 3. This method is in contrast to methods that rely on index barcodes for error correction.
[0165] I. Identification of Somatic Variants In various examples, enzymatically converted DNA is used to infer the methylation state of C residues in the genome. However, since enzymatic conversion of DNA converts unmethylated C residues to U residues and does not introduce other chemical changes to the DNA, somatic variants that do not correspond to C or T bases in the reference or query sequences can also be identified in the converted DNA. These somatic variants can be identified using existing methods designed for unconverted DNA, including error correction methods such as duplex sequencing. Furthermore, somatic variants corresponding to C or T bases in the reference or query sequences can be distinguished from methylation-related sequencing patterns using duplex sequencing based on the expectation that somatic variants should be found at the same position on both strands of the duplex DNA molecule, whereas methylation-related patterns should not (i.e., C and T bases are not seen base-paired with each other). This difference allows for the identification of both the methylation state of CpG sites and somatic variants in the EM sequences.
[0166] J. Estimation of Nucleosome Position Cytosine methylation at CpG sites can be highly enriched in DNA spanning nucleosomes compared to adjacent DNA. Thus, CpG methylation patterns may also be employed to infer the position of nucleosomes using machine learning approaches. EM sequence datasets may be analyzed following the same methods used for WGS to generate features that are input into machine learning methods and models, regardless of methylation conversion. Subsequently, the pattern of 5mC can be used to predict the position of nucleosomes, which can be useful for inference of gene expression and / or classification of diseases and cancers. In another example, features may be obtained from a combination of methylation state and nucleosome position information.
[0167] Metrics used in methylation analysis include, but are not limited to, M-bias (CpG, CHG, CHH base-wise methylation %), conversion efficiency (e.g., 100 average methylation % for CHH), hypomethylated blocks, methylation level (e.g., overall average methylation of CPG, CHH, CHG, chrM, LINE1, or ALU), dinucleotide coverage (normalized coverage of dinucleotides), equality of coverage (e.g., unique CpG sites at 1x and 10x average genomic coverage (when S4 is run), overall average CpG coverage (depth), and average coverage in CpG islands, CGI shelves, and CGI shores). In one example, the output of duplex-based CpG methylation calls is used as input for this analysis. In one example, information on fragment endpoints and lengths is used as feature input for the analysis. These metrics may be used as feature input for machine learning methods and models.
[0168] In another aspect, the present disclosure provides a method comprising: (a) providing a biological sample containing cfDNA from a subject; (b) exposing the cfDNA to conditions sufficient for any enrichment of methylated cfDNA in the sample; (c) enzymatically converting unmethylated cytosine nucleobases of the cfDNA to uracil nucleobases; (d) sequencing the cfDNA, thereby generating sequence reads; (e) (i) computer processing the sequence reads to determine the degree of methylation of the cfDNA based on the presence of uracil nucleobases and (ii) modeling at least partial degradation of the cfDNA, thereby generating degradation parameters; and (f) using the degradation parameters and the degree of methylation to determine genetic sequence features.
[0169] In some examples, sequencing the cfDNA includes determining the degree of methylation of the DNA based on the ratio of unconverted cytosine nucleobases to converted cytosine nucleobases. In some examples, the converted cytosine nucleobases are detected as uracil nucleobases. In some examples, the uracil nucleobases are observed as thymine nucleobases in the sequence reads.
[0170] In some examples, generating the degradation parameters includes using a Bayesian model. In some examples, the Bayesian model is based on strand bias or enzymatic conversion or over-conversion. In some examples, computer processing the sequence reads includes using the degradation parameters within the framework of a paired HMM or a naive Bayesian model.
[0171] K. Analysis of Variable Methylation Regions (DMRs) In one example, the methylation analysis is variable methylation region (DMR) analysis. DMRs are used to quantify CpG methylation across regions of the genome. The regions are dynamically assigned by discovery. By analyzing many samples of different classes, it is possible to identify the most variable methylation regions among various classifications. A subset of the regions can be selected to perform variable methylation and used for classification. The number of CpGs captured in the regions may also be used for analysis. In one example, the output of duplex-based CpG methylation calls is used as input for this analysis. The regions may be of variable size. In one example, a pre-discovery process is performed that bundles many CpG sites together as regions. In one example, DMRs are used as input features for machine learning methods and models.
[0172] L. Methylation Haplotype Block and Methylation Haplotype Load In one example, a haplotype block assay is applied to the sample. Identification of methylation haplotype blocks aids in the deconvolution of heterogeneous tissue samples and mapping of tumor tissue from plasma DNA. Tightly linked CpG sites, known as methylation haplotype blocks (MHBs), are identifiable in WGBS data. A metric called methylation haplotype load (MHL) is used to perform tissue-specific methylation analysis at the block level. This method provides blocks with a large amount of useful information for deconvolution of heterogeneous samples. This method is useful for quantitative estimation of tumor load and tissue origin mapping in circulating cfDNA. In one example, the output of duplex-based CpG methylation calls is used as input for this analysis. In one example, haplotype blocks are used as input features for machine learning methods and models.
[0173] M. Targeted Methylation Call Analysis for Identifying Cell-Type of Origin In one aspect, a method for targeted methylation calling is used to identify the cell type origin of cfDNA molecules based on methylation patterns. This method provides a probabilistic model of the co-methylation state of multiple adjacent CpG sites on individual sequencing reads in order to utilize the extensive nature of DNA methylation for signal amplification. This model develops the probability of sequencing reads for each cell type and then develops an overall cell type mixture model to fit the model.
[0174] In conventional DNA methylation analysis, attention has been paid to the methylation rate (β-value) of individual CpG sites in a cell population, indicating the proportion of cells in which that CpG site is methylated. Such population-averaged measurements are often not sensitive enough to capture abnormal methylation signals that only affect a portion of cfDNA. However, based on the extensive nature of DNA methylation, disease-specific cfDNA reads can be computationally distinguished from normal cfDNA reads.
[0175] In addition, considering the extensive nature of DNA methylation, the co-methylation state of multiple adjacent CpG sites can be used to easily distinguish cancer-specific cfDNA reads from normal cfDNA reads. The average value (referred to as the α-value) of the methylation values of all CpG sites in a given read provides a difference (0 and 1) between abnormal methylated cfDNA and normal cfDNA (αtumor = 0%, αnormal = 100%). The methylation α-value is used to estimate whether the binding probabilities of all CpG sites in the read follow the DNA methylation signature of the disease. This method can highly sensitively identify cfDNA derived from multiple cell types among all cfDNA in plasma.
[0176] In various examples, reads are aligned to a reference genome using an alignment tool, and methylated cytosines are called. PCR duplicates are removed, and the number of methylated and unmethylated cytosines is quantified for each CpG site. The methylation level of a CpG cluster is calculated as the ratio of the number of methylated cytosines to the total number of cytosines in the cluster. In this WGBS data processing procedure, the average methylation level of CpG clusters in normal plasma samples used for the identification of methylation markers is calculated. When a plasma cfDNA sample is used as test data, the combined methylation status of all CpG sites of individual sequencing reads aligned to the regions of the marker panel is extracted and input into a machine learning model. In this method, duplex-based CpG methylation calls are used as input features for the analysis of methylation status and the generation of features. To improve the input data quality of cfDNA methylation data with high coverage, reads covering CpG sites with 2-fold, 3-fold, or 4-fold or more can be filtered.
[0177] The methylation sequencing method described herein can improve the quality of sequencing reads, for example, by reducing PCR errors and biases and reducing the degradation of DNA that occurs during bisulfite conversion. In one example, methylation sequencing data is used to model overlapping regions. In one example, machine learning modeling can determine the cell type origin for the identified methylated DNA regions.
[0178] In various examples, the model can classify origins from two or more cell types. In other examples, the model can classify sequences into 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 50, 75, 100, or 100 or more different cell types.
[0179] N.DNA Hydroxymethylation Analysis In one aspect of the present invention, hydroxymethylation in the adapter nucleic acid is substituted in the adapter ligation step, and then, instead of using dioxygenase and βGT to bind 5mC and 5hmC, 5hmC sequencing is achievable by using only βGT to bind glucose to the 5hmC residues in the test nucleic acid library insert. When the resulting sequencing data is compared to a reference genome, all positions of C in the reference that correspond to C in the test sequence are interpreted as hydroxymethylated C, and all C in the reference that are shown as T in the test sequence are interpreted as unmodified C or methylated C. Thus, the data interpretation for hydroxymethylation analysis is the same as for methylation analysis.
[0180] In one aspect of the present invention, methylation and hydroxymethylation sequencing libraries can be compared to identify the level of each cytosine modification (e.g., 5m or 5mC) at single nucleotide resolution.
[0181] In one aspect of the present invention, since the readout of the hydroxymethylation status is the same as the methylation status, all analysis methods used for methylation sequencing data can be applied to hydroxymethylation sequencing data.
[0182] IV. COMPUTER SYSTEM AND MACHINE LEARNING METHOD A. SAMPLE CHARACTERISTICS As used herein, for purposes related to machine learning and pattern recognition, the term "feature" may refer to individual measurable properties or characteristics of an observed phenomenon. Features are usually numerical, although structural features such as character strings and graphs are used in syntactic pattern recognition. The concept of "feature" is related to the concept of explanatory variables used in statistical techniques such as linear regression.
[0183] In one embodiment, the features are input into a feature matrix for machine learning analysis.
[0184] For multiple assays, the system identifies a feature set as input to a machine learning model. The system performs an assay for each molecular class and forms a feature vector from the measurements. The system inputs the feature vector into a machine learning model and obtains an output classification as to whether a biological sample has a specified property.
[0185] In one embodiment, the machine learning model outputs a classifier that distinguishes between two groups or classes of individuals, i.e., features in a population of individuals or features of the population. In one embodiment, the classifier is a trained machine learning classifier.
[0186] In one embodiment, loci or features rich in biomarker information in cancer tissue are assayed to form a profile. A receiver operating characteristic (ROC) curve is useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing between two populations (e.g., individuals who respond to a therapeutic agent and those who do not). Typically, the feature data for the entire population (e.g., cases and controls) are sorted in ascending order based on the value of a single feature.
[0187] In some embodiments, the disease is advanced adenoma (AA), colorectal cancer (CRC), colorectal carcinoma, or inflammatory bowel disease.
[0188] The term "input feature" or "feature" refers to a variable used by a model to predict an output classification (label) of a sample, e.g., a condition, sequence content (e.g., a mutation), a proposed data collection operation, or a proposed treatment. The value of the variable is determinable for a sample and can be used to determine a classification. Examples of input features for genetic data include alignment variables related to the alignment of sequence data (e.g., sequence reads) to the genome, and unaligned variables, such as the sequence content of sequence reads, measurements of proteins or autoantibodies, or variables related to the average methylation level in a genomic region.
[0189] The value of a variable can be determined for a sample and can be used to determine a classification. Examples of input features of genetic data include aligned variables related to the alignment of array data (e.g., array reads) to the genome, and unaligned variables, such as the sequence content of array reads, measurements of proteins or autoantibodies, or variables related to the average methylation level in genomic regions. In various examples, genetic features such as, for example, V-plot measurements, FREE-C, cfDNA measurements of transcription start sites, and DNA methylation levels of cfDNA fragments are used as input features for machine learning methods and models.
[0190] In one example, the sequencing information includes information regarding a plurality of genetic features including, but not limited to, transcription start sites, transcription factor binding sites, open and closed states of chromatin, positions or occupancies of nucleosomes.
[0191] B. Data Analysis In some embodiments, the present disclosure provides a system, method, or kit having data analysis implemented in software applications, computing hardware, or both. In various embodiments, the analysis application or system includes at least a data receiving module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In one embodiment, the data receiving module can include a computer system that connects laboratory hardware or instruments to a computer system that processes laboratory data. In one embodiment, the data preprocessing module can include a hardware system or computer software that performs operations on the data as preparation for analysis. Examples of operations that can be applied to the data by the preprocessing module include affine transformation, noise removal operations, data cleaning, reformatting, or subsampling. The data analysis module can be specialized for the analysis of genomic data from one or more genomic materials, and can identify abnormal patterns related to diseases, pathologies, conditions, risks, conditions, or phenotypes, for example, by taking in assembled genomic sequences and performing probabilistic and statistical analyses. The data interpretation module can use analytical methods obtained from, for example, statistics, mathematics, or biology, to support the understanding of the relationship between the identified abnormal patterns and health status, functional status, prognosis, or risk. The data visualization module can use mathematical modeling, computer graphics, or rendering methods to create a visual representation of the data that can facilitate the understanding or interpretation of the results.
[0192] In various embodiments, machine learning methods are applied to distinguish samples in a population of samples. In one embodiment, the machine learning method is applied to identify samples between healthy samples and progressive adenoma samples.
[0193] In one embodiment, the one or more machine learning operations used to train a methylation-based prediction engine include generalized linear models, generalized additive models, nonparametric regression operations, random forest classifiers, spatial regression operations, Bayesian regression models, time series analysis, Bayesian networks, Gaussian networks, decision tree learning operations, artificial neural networks, recurrent neural networks, reinforcement learning operations, linear / nonlinear regression operations, support vector machines, clustering operations, and genetic algorithm operations.
[0194] In various embodiments, the computer processing method is selected from logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptrons (MLP), network clustering, neurofuzzy, and artificial neural networks.
[0195] In some embodiments, the methods disclosed herein can include computational analysis of nucleic acid sequence data from a sample from an individual or from multiple individuals. The analysis can identify variants inferred from the sequence data based on probabilistic modeling, statistical modeling, machine modeling, network modeling, or statistical inference. Non-limiting examples of analysis methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline mutations or somatic mutations. In some embodiments, the variant can refer to a known variant. A known variant can be scientifically confirmed or reported in the literature. In some embodiments, the variant can refer to a hypothesized variant associated with a biological change. The biological change can be known or unknown. In some embodiments, the hypothesized variant may be reported in the literature but not yet biologically confirmed.
[0196] Alternatively, a putative variant may not be reported in the literature but can be inferred based on the computer analysis disclosed herein. In some embodiments, a germline variant can refer to a nucleic acid that causes a natural or normal variation.
[0197] Natural or normal variations can include, for example, skin color, hair color, and standard body weight. In some embodiments, a somatic mutation can refer to a nucleic acid that causes an acquired or abnormal mutation. Acquired or abnormal mutations can include, for example, cancer, obesity, diseases, symptoms, disorders, and impairments. In some embodiments, the analysis can include distinguishing germline variants. Germline variants can include, for example, private variants and somatic mutations. In some embodiments, the identified variants can be used by a clinician or other healthcare professional to improve healthcare methodologies, diagnostic accuracy, and cost reduction.
[0198] Furthermore, improved methods and computing systems or software media that can distinguish amplification and / or sequencing techniques, somatic mutations, and sequence errors of nucleic acids introduced by germline variants are provided herein. The methods provided may include simultaneously calling and scoring variants from aligned sequencing data of all samples obtained from a patient.
[0199] Samples obtained from subjects other than patients can also be used. Other samples can also be collected from subjects that have been previously analyzed by a sequencing assay or a targeted sequencing assay (i.e., a target resequencing assay). The methods, computing systems, or software media disclosed herein can improve the identification and accuracy of variants or mutations (e.g., germline or somatic, including copy number polymorphisms, single nucleotide polymorphisms, indels, gene fusions), and can reduce the detection limit by reducing the number of false positive and false negative identifications.
[0200] C. Classifier Generation In one aspect, the system and method provide a classifier generated based on feature information derived from the analysis of methylated sequences from a biological sample of cfDNA. The classifier forms part of a prediction engine for distinguishing groups within a population based on methylated sequence features identified in biological samples such as cfDNA.
[0201] In one embodiment, the classifier formats similar portions of methylation information into a unified format and a unified scale, stores the normalized methylation information in a column-oriented database, and trains a methylation prediction engine by applying one or more machine learning operations to the stored normalized methylation information, wherein the methylation prediction engine maps a combination of one or more features for a particular population, applies the methylation prediction engine to the accessed field information to identify methylation associated with a group, and classifies the individual into one group.
[0202] Specificity can be defined as the probability that a test among people without the disease is negative. Specificity is equal to the number of people without the disease who are determined to be negative divided by the total number of individuals without the disease.
[0203] In various embodiments, the model, classifier, or prediction test has a specificity of at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.
[0204] Sensitivity can be defined as the probability that a test among people without the disease is positive. Sensitivity is equal to the number of individuals with the disease who are determined to be negative divided by the total number of individuals with the disease.
[0205] In various embodiments, the model, classifier, or prediction test has a sensitivity of at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.
[0206] In one embodiment, the group is selected from healthy (asymptomatic) inflammatory bowel disease, AA, or CRC.
[0207] D. Digital Processing Device In some embodiments, the subject matter described herein may include a digital processing device or its use. In some embodiments, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that execute the functions of the device. In some embodiments, the digital processing device may include an operating system configured to execute executable instructions. In some embodiments, the digital processing device may be optionally connected to a computer network. In some embodiments, the digital processing device may be optionally connected to the Internet so as to connect to the World Wide Web. In some embodiments, the digital processing device may be optionally connected to a cloud computing infrastructure. In some embodiments, the digital processing device may be optionally connected to an intranet. In some embodiments, the digital processing device may be optionally connected to a data storage device.
[0208] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, handheld computers, Internet appliances, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, booklets, slates, and convertible configurations.
[0209] In some embodiments, the digital processing device may include an operating system configured to execute executable instructions. For example, the operating system may include software that includes programs and data, and the software manages the hardware of the device and provides services for the execution of applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Non-limiting examples of suitable personal computer operating systems include Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX®-based operating systems, such as GNU / Linux®. In some embodiments, the operating system may be provided by cloud computing, and the cloud computing resources may be provided by one or more service providers.
[0210] In some embodiments, the device may include a storage device and / or a memory device. The storage device and / or the memory device can be one or more physical devices used to store data or programs, either temporarily or permanently. In some embodiments, the device can be a volatile memory and requires power to maintain the stored information. In some embodiments, the device can be a non-volatile memory and can retain the stored information when power is not supplied to the digital processing device. In some embodiments, the non-volatile memory can include flash memory. In some embodiments, the non-volatile memory can include dynamic random access memory (DRAM). In some embodiments, the non-volatile memory can include ferroelectric random access memory (FRAM (registered trademark)). In some embodiments, the non-volatile memory can include phase change random access memory (PRAM). In some embodiments, the device can be a storage device, including, for example, CD-ROM, DVD, flash memory device, magnetic disk drive, magnetic tape drive, optical disk drive, and cloud computing-based storage devices. In some embodiments, the storage device and / or the memory device can be a combination of devices such as those disclosed herein.
[0211] In some embodiments, the digital processing device may include a display for sending visual information to the user. In some embodiments, the display may be a cathode ray tube (CRT). In some embodiments, the display may be a liquid crystal display (LCD). In some embodiments, the display may be a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display may be an organic light emitting diode (OLED) display. In some embodiments, the OLED display may be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some embodiments, the display may be a plasma display. In some embodiments, the display may be a video projector. In some embodiments, the display may be a combination of devices such as those disclosed herein.
[0212] In some embodiments, the digital processing device may include an input device for receiving information from the user. In some embodiments, the input device may be a keyboard. In some embodiments, the input device may be a pointing device, including, for example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some embodiments, the input device may be a touch screen or a multi-touch screen. In some embodiments, the input device may be a microphone for capturing voice or other audio input. In some embodiments, the input device may be a video camera for capturing motion or visual input. In some embodiments, the input device may be a combination of devices such as those disclosed herein.
[0213] E. Non-transitory computer-readable storage medium In some embodiments, the subject matter disclosed herein may include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of a digital processing device optionally connected to a network. In some embodiments, the computer-readable storage media may be a tangible component of the digital processing device. In some embodiments, the computer-readable storage media may be optionally removable from the digital processing device. In some embodiments, the computer-readable storage media may include, for example, CD-ROMs, DVDs, flash memory devices, solid state memories, magnetic disk devices, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In some embodiments, the program and instructions may be encoded on the media permanently, nearly permanently, semi-permanently, or non-transitorily.
[0214] F. Computer System The present disclosure provides a computer system programmed to implement the methods of the present disclosure. FIG. 5 shows a computer system (501) programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, or reference sequences. The computer system (501) can process various aspects of the patient data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system (501) can be an electronic device of a user or computer system located remotely from the electronic device. The electronic device may be a mobile electronic device.
[0215] A computer system (501) includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") (505), which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (501) includes a memory or storage location (510) (e.g., random access memory, read-only memory, flash memory), an electronic storage device (515) (e.g., hard disk), a communication interface (520) (e.g., network adapter) for communicating with one or more other systems, and peripheral devices (525), such as cache, other memory, data storage devices, and / or an electronic display adapter. The memory (510), storage device (515), interface (520), and peripheral devices (525) communicate with the CPU (505) via a communication bus (solid line), such as a motherboard. The storage device (515) can be a data storage device (or data repository) for storing data. The computer system (501) can be operably connected to a computer network ("network") (530) with the help of the communication interface (520). The network (530) can be the Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some embodiments, the network (530) is a telecommunications and / or data network. The network (530) can include one or more computer servers, which can enable distributed computing such as cloud computing. In some embodiments, the network (530) can implement a peer-to-peer network with the help of the computer system (501), thereby enabling devices connected to the computer system (501) to act as clients or servers.
[0216] The CPU (505) can execute a series of machine-readable instructions, which can be embodied in a program or software. These instructions can be stored in a storage location such as the memory (510). These instructions can be directed to the CPU (505), which can later program or otherwise configure the CPU (505) to implement the method of the present disclosure. Examples of operations performed by the CPU (505) include fetch, decode, execute, and write-back.
[0217] The CPU (505) can be part of a circuit such as an integrated circuit. One or more other components of the system (501) may be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).
[0218] The storage device (515) can store files such as drivers, libraries, and saved programs. The storage device (515) can store user data, e.g., user preferences and user programs. The computer system (501) can include one or more additional data storage devices outside of the computer system (501), such as being located on a remote server in communication with the computer system (501) via an intranet or the Internet in some embodiments.
[0219] A computer system (501) can communicate with one or more remote computer systems via a network (530). For example, the computer (501) can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or personal digital assistants. A user can access the computer system (501) via the network (530).
[0220] The methods described herein can be executed by machine (e.g., computer processor) executable code stored in an electronic memory location of the computer system (501), such as, for example, on a memory (510) or an electronic storage device (515). The machine executable code or machine readable code can be provided in the form of software. In use, the code can be executed by a processor (505). In some embodiments, the code can be retrieved from the storage device (515) and stored in the memory (510) for immediate access by the processor (505). In some embodiments, the electronic storage device (515) may be excluded and the machine executable instructions may be stored in the memory (510).
[0221] The code can be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or can be interpreted or compiled during run-time. The code can be provided in a programming language selected to enable the code to be executed in a pre-compiled, interpreted, or as-compiled fashion.
[0222] Aspects of the systems and methods provided herein, such as computer system (501), may be embodied in programming. Various aspects of this technology may typically be considered as a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data that is executed or embodied on a type of machine-readable medium. The machine executable code can be stored in an electronic storage device such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of various semiconductor memories, tape drives, disk drives, etc., which are tangible memories of a computer or processor, or associated modules thereof, and which can provide a non-transitory recording medium for software programming at any time. All or part of the software is sometimes communicated via the Internet or various other electrical communication networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves such as those used over physical interfaces between local devices via wired and optical terrestrial communication line networks and over various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless restricted to non-transitory and tangible "storage" media, terms such as computer or machine "readable media" refer to media involved in providing instructions to a processor for execution.
[0223] Accordingly, machine-readable media such as computer-executable code may take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, any of the storage devices in a computer such as optical disks or magnetic disks, such as those used to implement a database shown in the drawings. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes copper wire and fiber optics including coaxial cable and wires that form a bus within a computer system. Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Accordingly, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROM, DVD, or DVD-ROM, other optical media, punch cards, paper tape, other physical storage media with patterns of holes, RAM, ROM, PROM, and EPROM, FLASH®-EPROM, other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0224] A computer system (501) may include or be in communication with an electronic display (135) including, for example, a user interface (UI) (540) for providing nucleic acid sequences, concentrated nucleic acid samples, expression profiles, and analysis of expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0225] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit (505). The algorithms can, for example, search for a plurality of adjustment elements, sequence a nucleic acid sample, concentrate a nucleic acid sample, determine an expression profile of a nucleic acid sample, analyze the expression profile of the nucleic acid sample, and record or disseminate the results of the analysis of the expression profile.
[0226] In some embodiments, the subject matter disclosed herein includes at least one computer program, or the use of such computer program. The computer program can be executable on a CPU, GPU, or TPU of a digital processing device and can be a series of instructions written to perform a particular task. The computer-readable instructions can be executed as program modules such as functions, objects, application programming interfaces (APIs), data structures, etc. that perform a particular task or execute a particular extracted data type. Considering the disclosure provided herein, one of ordinary skill in the art will recognize that the computer program can be written in various versions of various languages.
[0227] The functionality of the computer-readable instructions can be combined or distributed according to the requirements of various environments. In some embodiments, the computer program may include a single sequence of instructions. In some embodiments, the computer program may include multiple sequences of instructions. In some embodiments, the computer program may be provided from a single location. In some embodiments, the computer program may be provided from multiple locations. In some embodiments, the computer program may include one or more software modules. In some embodiments, the computer program may include, in whole or in part, one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plugins, extensions, add-ons, or add-ins, or combinations thereof.
[0228] In some embodiments, the computer processing can be in the form of methods of statistics, mathematics, biology, or any combination thereof. In some embodiments, the computer processing method can include, for example, dimensionality reduction methods such as logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, network clustering, neural networks.
[0229] In some embodiments, the computer processing method is a supervised machine learning method that includes, for example, regression, support vector machines, tree-based methods, and networks.
[0230] In some embodiments, the computer processing method is an unsupervised machine learning method that includes, for example, clustering, networks, principal component analysis, and matrix factorization.
[0231] G. Database In some embodiments, the subject matter disclosed herein includes one or more databases, or the use thereof, for storing patient data, biological data, biological sequences, or reference sequences. The reference sequences can be derived from the database. Given the disclosure provided herein, one of ordinary skill in the art will recognize that many databases are suitable for storing and retrieving sequence information. In some embodiments, suitable databases can include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some embodiments, the database can be Internet-based. In some embodiments, the database can be web-based. In some embodiments, the database can be cloud-computing-based. In some embodiments, the database can be based on one or more local computer storage devices.
[0232] V. Cancer Diagnosis and Detection The trained machine learning methods, models, and discriminative classifiers described herein are useful for a variety of medical applications including cancer detection, diagnosis, and treatment responsiveness. When the model is trained using features from individual metadata and analysis, its use can be adapted to stratify individuals in a population and accordingly guide treatment decisions.
[0233] A. Diagnosis The methods and systems provided herein can perform predictive analysis using an artificial intelligence-based approach to analyze data obtained from a subject (patient) in order to generate a diagnostic output for a subject having cancer (e.g., CRC). For example, to generate a diagnosis of a subject having cancer, the use can apply a predictive algorithm to the data obtained. The predictive algorithm can include an artificial intelligence-based predictor such as a machine learning-based predictor configured to process the data obtained to generate a diagnosis of a subject having cancer.
[0234] A machine learning predictor can be trained using a dataset obtained from one or more sets of a cohort of cancer patients as input to the machine learning predictor and the results of known diagnoses (e.g., stage classification and / or tumor proportion) of the subjects as output, for example, a dataset generated by performing a multi - analysis assay on a biological sample of an individual.
[0235] Training of the dataset (e.g., a dataset generated by performing a multi - analysis assay on a biological sample of an individual) can be generated, for example, from one or more sets of subjects having common characteristics (features) and outcomes (labels). Training of the dataset can include a set of features and labels corresponding to features related to the diagnosis. The features can include, for example, certain ranges or categories of cfDNA assay measurements, such as the number of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap or fall within each of a set of bins (genomic windows) of a reference genome. For example, a set of features collected from a given subject at a given time point can function collectively as a diagnostic signature and can indicate the identified cancer of the subject at that given time. The characteristics can also include labels indicating the diagnostic results of the subject, such as for one or more cancers.
[0236] The labels can include, for example, outcomes such as the results of known diagnoses (e.g., stage classification and / or tumor proportion) of the subjects. The outcomes can include characteristics related to cancer in the subject. For example, the characteristics can indicate that the subject has one or more cancers.
[0237] A training set (e.g., a training data set) can be selected by random extraction of a set of data corresponding to one or more subjects (e.g., a retrospective cohort and / or a prospective cohort of patients with or without one or more cancers). Alternatively, a training set (e.g., a training data set) can be selected by proportional extraction of a set of data corresponding to one or more subjects (e.g., a retrospective cohort and / or a prospective cohort of patients with or without one or more cancers). The training set can be balanced across multiple sets of data corresponding to one or more sets of subjects (e.g., patients from various clinical facilities or clinical trials). The machine learning predictor may be trained until predefined conditions for accuracy or performance are met, such as having a minimum target value corresponding to a measure of diagnostic accuracy. For example, the measure of diagnostic accuracy can correspond to the diagnosis, staging, or prediction of the proportion of tumors of one or more cancers in a subject.
[0238] Examples of measures of diagnostic accuracy can include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC) of the ROC curve corresponding to the diagnostic accuracy of detecting or predicting cancer (e.g., colorectal cancer).
[0239] In another aspect, the present disclosure provides a method for identifying cancer in a subject, the method comprising: (a) providing a biological sample comprising cell-free nucleic acid (cfNA) molecules from the subject; (b) performing methylation sequencing on the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads in each of a first plurality of genomic regions of the reference genome to generate a first cfNA signature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 different regions (each of the at least about 10 different regions); and (e) applying a trained algorithm to the first cfNA signature set to generate a likelihood that the subject has cancer.
[0240] For example, such predefined conditions can include a sensitivity for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0241] As another example, such predefined conditions can include a specificity for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0242] As another example, such a predefined condition may be that the positive predictive value (PPV) for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0243] As another example, such a predefined condition may be that the negative predictive value (NPV) for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0244] As another example, such a predefined condition may be that the area under the ROC curve (AUC) for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) is at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0245] In some examples of any of the foregoing aspects, the method further includes a step of monitoring the progression of the disease in a subject, where the monitoring step is at least partially based on the characteristics of the gene sequence. In some examples, the disease is cancer.
[0246] In some examples of any of the foregoing aspects, the method further comprises determining the tissue origin of cancer in a subject, wherein the determining step is at least partially based on characteristics of a gene sequence.
[0247] In some examples of any of the foregoing aspects, the method further comprises estimating the tumor burden of a subject, wherein the estimating step is at least partially based on characteristics of a gene sequence.
[0248] B. Treatment Responsiveness The predictive classifiers, systems, and methods described herein are useful for classifying populations of individuals for many clinical applications (e.g., based on performing multi-analyte assays on biological samples of an individual). Examples of such clinical applications include detecting early-stage cancer, diagnosing cancer, classifying cancer at a particular stage of a disease, or determining responsiveness or resistance to therapeutic agents for treating cancer.
[0249] The methods and systems described herein are applicable to various cancer types as well as grades and stages, and are therefore not limited to a single cancer disease type. Thus, the combination of analyses and assays can be used in the present systems and methods to predict the responsiveness of cancer therapies across various cancer types in various tissues and to classify individuals based on treatment responsiveness. In one example, the classifiers described herein stratify a group of individuals into treatment responders and non-responders.
[0250] The present disclosure further provides a method for determining a drug target (e.g., a gene associated with / important to a particular class) of a disease or disorder in a subject, the method comprising evaluating a sample obtained from the individual for the gene expression level of at least one gene, and using a proximity analysis routine to determine a gene associated with the classification of the sample, thereby identifying one or more drug targets associated with the classification.
[0251] The present disclosure further provides a method for determining the effectiveness of a drug designed to treat a disease class, the method comprising obtaining a sample from an individual having the disease class, exposing the sample to the drug, evaluating the sample exposed to the drug with respect to the gene expression level of at least one gene, and using a computer model constructed using a weighted voting scheme to classify the sample exposed to the drug into the disease class according to the relative gene expression level of the sample with respect to the relative gene expression level of the model.
[0252] The present disclosure further provides a method for determining the effectiveness of a drug designed to treat a disease class, wherein the individual has been exposed to the drug, the method comprising obtaining a sample from the individual exposed to the drug, evaluating the sample with respect to the gene expression level of at least one gene, and using a model constructed using a weighted voting scheme to classify the sample into the disease class, the use comprising evaluating the gene expression level of the sample compared to the gene expression level of the model.
[0253] However, another application is a method for determining whether an individual belongs to a phenotypic class (e.g., intelligence, response to treatment, longevity, susceptibility to viral infection, or obesity), the method comprising obtaining a sample from the individual, evaluating the sample with respect to the gene expression level of at least one gene, and using a model constructed using a weighted voting scheme to classify the sample into a disease class, the use comprising evaluating the gene expression level of the sample compared to the gene expression level of the model.
[0254] Biomarkers may be useful for predicting the prognosis of colorectal cancer patients. The ability to classify patients as high risk (poor prognosis) or low risk (good prognosis) allows for the selection of appropriate treatments for these patients. For example, high-risk patients may benefit from aggressive treatment, while for low-risk patients, treatment may have no significant benefit.
[0255] There are predictive biomarkers that can guide treatment decisions by identifying subsets of patients who may be "exceptional responders" to a particular cancer treatment or subsets of individuals who may benefit from alternative therapies.
[0256] In one aspect, the systems and methods described herein for classifying populations based on treatment responsiveness refer to, but are not limited to, cancers treated by chemotherapeutic agents of the class DNA-damaging agents, DNA repair target therapies, inhibitors of DNA damage signaling, inhibitors of DNA damage-induced cell cycle arrest, and inhibition of processes that indirectly lead to DNA damage. Each of these chemotherapeutic agents can be considered a "DNA damage therapeutic agent".
[0257] Patient analyte data is classified into high-risk and low-risk patient groups, such as patients with a high or low risk of clinical recurrence, and the results can be used to determine treatment strategies. For example, a patient determined to be a high-risk patient may be treated with adjuvant chemotherapy after surgery. In the case of a patient considered to be a low-risk patient, adjuvant chemotherapy may be withheld after surgery. Accordingly, the present disclosure provides, in one aspect, a method for preparing a gene expression profile of a colorectal cancer tumor indicative of recurrence risk.
[0258] In various examples, the classifiers described herein stratify a population of individuals between responders and non-responders to treatment.
[0259] In various examples, the treatment is selected from alkylating agents, plant alkaloids, antitumor antibiotics, antimetabolites, topoisomerase inhibitors, retinoids, checkpoint inhibitor therapy, and VEGF inhibitors.
[0260] Examples of treatments by which a population can be stratified into responders and non-responders include, but are not limited to, sorafenib, regorafenib, imatinib, eribulin, gemcitabine, capecitabine, pazopanib, lapatinib, dabrafenib, sunitinib malate, crizotinib, everolimus, torisirolimus, sirolimus, axitinib, gefitinib, anastrozole, bicalutamide, fulvestrant, ralitrexed, pemetrexed, goserelin acetate, erlotinib, bemrafenib, besmudegebu, tamoxifen citrate, paclitaxel, docetaxel, cabazitaxel, oxaliplatin, ziv-aflibercept, bevacizumab, trastuzumab, pertuzumab, panitumumab, taxane, bleomycin, melphalen, plumbagin, camptosar, mitomycin C, mitoxantrone, poly(styrene maleic acid)-conjugated neocarzinostatin (SMANCS), doxorubicin, pegylated doxorubicin, FOLFORI, 5-fluorouracil temozolomide, pasireotide, tegafur, gimeracil, oteraci, itraconazole, bortezomib, lenalidomide, irinotecan, epirubicin, romidepsin, resminostat, tasquinimod, refametinib, lapatinib, Tyverb (registered trademark), Arenegyr, NGR-TNF, pasireotide, Signifor (registered trademark), ticilimumab, tremelimumab, lansoprazole, PrevOnco (registered trademark), ABT-869, linifanib, vorolanib, cibantinib, Tarceva (registered trademark), erlotinib, Stivarga (registered trademark), regorafenib, fluoro-sorafenib, brivanib, liposomal doxorubicin, lenvatinib, ramucirumab, peretinoin, Ruchiko, muparfostat, Teysuno (registered trademark), tegafur, gimeracil, oteracil, and orantinib, chemotherapeutic agents, and antibody therapies including alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, or rituximab.
[0261] In other examples, the population may be stratified into responders and non-responders to checkpoint inhibitor therapy, such as compounds that bind to PD-1 or CTLA4.
[0262] In other examples, the population may be stratified into responders and non-responders to anti-VEGF therapy that binds to VEGF pathway targets.
[0263] VI. Indications In some examples, the biological state may include a disease. In some examples, the biological state may be a disease stage. In some examples, the biological state may be a gradual change in the physiological state. In some examples, the biological state may be a treatment effect. In some examples, the biological state may be a drug effect. In some examples, the biological state may be a surgical effect. In some examples, the biological state may be the physiological state after lifestyle improvement. Non-limiting examples of lifestyle improvement include changes in diet, changes in smoking, and changes in sleep patterns. In some examples, the biological state is unknown. The analysis described herein may include machine learning to infer or interpret an unknown biological state.
[0264] In one example, the present system and method are particularly useful for applications related to colon cancer, i.e., cancer that occurs in the tissues of the colon, the longest part of the large intestine. Most colon cancers are adenocarcinomas (cancers that start in cells that form gland-like structures lining internal organs). The progression of cancer is characterized by the stage or extent of the cancer in the body. Staging classifications are typically based on the size of the tumor, whether lymph nodes contain cancer, and whether the cancer has spread from the site where it first occurred to other parts of the body. The stages of colon cancer include stage I, stage II, stage III, and stage IV. Unless otherwise specified, the term "colon cancer" refers to colon cancer at stage 0, stage I, stage II (including stage IIA or IIB), stage III (including stage IIIA, IIIB, or IIIC), or stage IV. In some of the examples described herein, the colon cancer can be of any stage. In some examples, the colon cancer is stage I colon cancer. In some examples, the colon cancer is stage II colon cancer. In some examples, the colon cancer is stage III colon cancer. In some examples, the colon cancer is stage IV colon cancer.
[0265] Diseases that can be inferred by the disclosed method include, for example, cancer, intestinal-related diseases, immune-mediated inflammatory diseases, nervous system diseases, kidney diseases, prenatal diseases, and metabolic diseases.
[0266] In some examples, the methods of the present disclosure can be used to diagnose cancer. Non-limiting examples of cancer include adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0267] Non-limiting examples of cancers that can be inferred by the disclosed methods and systems include acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, cholangiocarcinoma, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brainstem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumor, breast cancer, bronchial tumor, Burkitt lymphoma, non-Hodgkin lymphoma, carcinoid tumor, cervical cancer, chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrous histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell cancer, melanoma, oral cancer, myelodysplastic syndrome, multiple myeloma, medulloblastoma, nasal cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papilloma, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumor, plasma cell tumor, prostate cancer, rectal cancer, renal cell cancer, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenström macroglobulinemia, and Wilms tumor.
[0268] Non-limiting examples of bowel-related diseases that can be inferred by the disclosed methods and systems include Crohn's disease, colitis, ulcerative colitis (UC), inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), and celiac disease. In some examples, the disease is inflammatory bowel disease, colitis, ulcerative colitis, Crohn's disease, microscopic colitis, collagenous colitis, lymphocytic colitis, diversion colitis, Behçet's disease, and ulcerative colitis.
[0269] Non-limiting examples of immune-mediated inflammatory diseases that can be inferred by the disclosed methods and systems include psoriasis, sarcoidosis, rheumatoid arthritis, asthma, rhinitis (hay fever), food allergies, eczema, lupus, multiple sclerosis, fibromyalgia, type 1 diabetes, and Lyme disease. Non-limiting examples of neurological diseases that can be inferred by the disclosed methods and systems include Parkinson's disease, Huntington's disease, multiple sclerosis, Alzheimer's disease, stroke, epilepsy, neurodegeneration, and neuropathy. Non-limiting examples of kidney diseases that can be inferred by the disclosed methods and systems include interstitial nephritis, acute kidney injury, and nephrosis. Non-limiting examples of prenatal diseases that can be inferred by the disclosed methods and systems include Down syndrome, aneuploidy, spina bifida, trisomy, Edwards syndrome, teratoma, sacrococcygeal teratoma (SCT), ventricular enlargement, renal agenesis, cystic fibrosis, and fetal hydrops. Non-limiting examples of metabolic diseases that can be inferred by the disclosed methods and systems include cystinosis, Fabry disease, Gaucher disease, Lesch-Nyhan syndrome, Niemann-Pick disease, phenylketonuria, Pompe disease, Tay-Sachs disease.
[0270] The specific details of the specific examples may be combined in any suitable manner without departing from the spirit and scope of the disclosed examples of the present invention. However, other examples of the present invention may be directed to specific examples of the individual aspects or specific combinations of these individual aspects. All patents, patent applications, publications, and descriptions mentioned herein are hereby incorporated by reference in their entirety for all purposes.
[0271] VII. Kit The present disclosure provides a kit for identifying or monitoring cancer in a subject. The kit includes probes for identifying a quantitative measure (e.g., indicating presence, absence, or relative amount) of sequences at each of a plurality of cancer-related genomic loci in a cell-free biological sample of the subject. A quantitative measure (e.g., indicating presence, absence, or relative amount) of sequences at each of the plurality of cancer-related genomic loci in the cell-free biological sample can indicate one or more cancers. The probes can be selective for sequences at a plurality of cancer-related genomic loci in the cell-free biological sample. The kit includes instructions for using the probes to process the cell-free biological sample to generate a dataset indicating a quantitative measure (e.g., indicating presence, absence, or relative amount) of sequences at each of the plurality of cancer-related genomic loci in the cell-free biological sample of the subject. In one embodiment, the kit includes a primer set, PCR reaction components, sequencing reagents, minimally disruptive conversion reagents, and library preparation reagents.
[0272] The probes in the kit can be selective for sequences at a plurality of cancer-related genomic loci in the cell-free biological sample. The probes in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to the plurality of cancer-related genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can have sequence complementarity with nucleic acid sequences from one or more of the plurality of cancer-related genomic loci or genomic regions. The plurality of cancer-related genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more different cancer-related genomic loci or genomic regions identified for targeted methylation sequencing.
[0273] The instructions in the kit include instructions for analyzing a cell-free biological sample using probes selective for sequences at multiple cancer-related genomic loci in the cell-free biological sample. These probes can be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) from one or more of the multiple cancer-related genomic loci. These nucleic acid molecules can be primers or enrichment sequences. The instructions for analyzing the cell-free biological sample include performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample in order to generate a dataset indicative of a quantitative measure of the sequences at each of the multiple cancer-related genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative amount). A quantitative measure of the sequences at each of the multiple cancer-related genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative amount) can indicate one or more cancers.
[0274] The instructions in the kit include instructions for measuring and interpreting assay readouts, which can be quantified at one or more of the multiple cancer-related genomic loci in order to generate a dataset indicative of a quantitative measure of the sequences at each of the multiple cancer-related genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative amount). For example, quantifying array hybridization or polymerase chain reaction (PCR) corresponding to the multiple cancer-related genomic loci can generate a dataset indicative of a quantitative measure of the sequences at each of the multiple cancer-related genomic loci in the cell-free biological sample (e.g., indicating presence, absence, or relative amount). Assay readouts can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized versions thereof.
Example
[0275] Example 1: Preparation of a targeted EM-seq library and generation of a classifier.
[0276] Starting material: 10 - 200 ng of double-stranded DNA.
[0277] 1. DNA preparation EDTA was removed from the DNA before oxidation, and the DNA sample had a final volume of 29 μl. Control DNA was used to evaluate oxidation and deamination. For sequencing on the Illumina platform, refer to the Enzymatic Methyl-seq Kit Manual (NEB #E7120) for usage recommendations.
[0278] 2. Adapter ligation 3. Oxidation of 5-methylcytosine and 5-hydroxymethylcytosine TET2 buffer was prepared. Next, the TET2 reaction was added to one tube containing the Reaction Buffer Supplement, and then mixed thoroughly. On ice, the TET2 reaction buffer, oxidation supplement, oxidation accelerator, and TET2 enzyme were added directly to the DNA sample. Then, the mixture was mixed thoroughly by vortexing. After a short spin in a centrifuge, the diluted Fe(II) solution was added to the mixture. Then, the mixture was mixed thoroughly by vortexing or by pipetting up and down and spun in a centrifuge for a short time. Next, the mixture was incubated in a thermocycler at 37 °C for 1 hour. Then, the sample was transferred to ice and treated with 1 μl of stop reagent (yellow). Next, the mixture was mixed thoroughly by vortexing or by pipetting up and down at least 10 times and spun in a centrifuge for a short time. Finally, the mixture was incubated in a thermocycler at 37 °C for 30 minutes and then at 4 °C.
[0279] 4. Cleanup of TET2-converted DNA The sample purification beads were resuspended by vortexing. Next, NEBNext sample purification beads were added to each sample and then thoroughly mixed by pipetting up and down. The samples were incubated on the bench top at room temperature for at least 5 minutes. Subsequently, multiple tubes were placed upright on an appropriate magnetic stand to separate the beads from the supernatant. After 5 minutes (or when the solution became clear), the supernatant was carefully removed and discarded without disturbing the beads containing the DNA target. While remaining on the magnetic stand, freshly prepared 80% ethanol was added to each of the multiple tubes. The samples were incubated at room temperature for 30 seconds and then the supernatant was carefully removed and discarded. The washing was repeated once for a total of 2 washes. After the second wash using a p10 pipette tip, all visible liquid was removed. Then, with the lids open, the tubes were left on the magnetic stand and the beads were air-dried for 2 minutes. Thereafter, the tubes were removed from the magnetic stand. DNA was eluted from the beads using elution buffer. The elution buffer was added to each tube and thoroughly mixed by pipetting up and down 10 times. Subsequently, the samples were incubated at room temperature for at least 1 minute. If necessary, before returning the tubes to the magnetic stand, the samples were briefly centrifuged to collect the liquid from the sides of the tubes. Then, the tubes were returned to the magnetic stand. After 3 minutes (or whenever the solution became clear), the eluted DNA from the supernatant was transferred to a new PCR tube.
[0280] 5. Denaturation of DNA Prior to the deamination of cytosine, the DNA was denatured using either formamide or 0.1 N sodium hydroxide.
[0281] 6. Deamination of cytosine On ice, APOBEC reaction buffer, BSA, and APOBEC were added to the denatured DNA. Subsequently, before briefly centrifuging, the mixture was thoroughly mixed by vortexing or by pipetting up and down at least 10 times. Then, the above mixture was incubated in a thermocycler at 37 °C for 3 hours and then at 4 °C.
[0282] 7. Cleanup of Deaminated DNA The sample purification beads were resuspended by vortexing. Next, 100 μl of resuspended NEBNext sample purification beads were added to each sample and then mixed thoroughly by pipetting up and down at least 10 times. During the final mixing, all the liquid was carefully expelled from the tip. Then, the samples were incubated on the bench top at room temperature for at least 5 minutes. After 5 minutes (or when the solution became clear), the supernatant was carefully removed and discarded. While placed on the magnetic stand, freshly prepared 80% ethanol was added to the above tubes. Then, the samples were incubated at room temperature for 30 seconds and then the supernatant was carefully removed and discarded. The wash was repeated once for a total of 2 washes. Next, with the lids open, the tubes were left on the magnetic stand and the beads were air-dried for 90 seconds. Then, the DNA target was eluted from the beads using elution buffer. The elution buffer was added to each of the tubes and mixed thoroughly by pipetting up and down 10 times. The samples were incubated at room temperature for at least 1 minute. If necessary, the samples were briefly centrifuged to collect the liquid from the sides of the tubes before returning the tubes to the magnetic stand. Then, the tubes were returned to the magnetic stand. After 3 minutes (or whenever the solution became clear), the eluted DNA target in the supernatant was transferred to a new PCR tube.
[0283] 8. Multiplex Amplification and Targeted Methylation Classification To enable targeted methylation analysis of pre-identified regions of the genome, raw data files were used for alignment and methylation calling using conventional tools. Whole genome amplification of enzymatically converted DNA was performed. Target enrichment was carried out on the enzymatically converted library, and 5'-biotinylated capture probes were used to specifically pull down pre-identified DNA fragments containing the target CpG sites. Hybrid selection was performed using the Illumina TruSightVR Rapid Capture Kit. In the hybridization step, Capture Target Buffer 3 (Illumina) was used instead of the enrichment hybridization buffer. After hybridization, the captured DNA fragments were amplified with 14 PCR cycles. The target capture library was sequenced on an Illumina HiSeqVR 2500 sequencer using a 2x100 cycle run with 4 - 5 samples in rapid run mode. The base diversity was increased and the sequencing quality was improved by spiking 10% PhiX into the enzymatic sequencing library.
[0284] Using the conventional method, FASTQ files were mapped to the reference genome, and methylation scores were calculated for disease classification. Feature data containing a set of CpG sites related to health, disease, disease status, and treatment responsiveness were input into a machine learning model to identify classifiers for stratifying individuals in the population.
[0285] Example 2: Targeted EM-seq using a conversion-resistant sequencing adapter / primer system Identified adapters of a known array were ligated to the ends of DNA molecules in a sample having an unknown array. Subsequently, the adapters were used to PCR amplify a whole library of various molecules using a single set of primers corresponding to the known adapters. During subsequent sequencing reactions, the ligated adapter sequences were further used as binding sites for sequencing primers. To utilize data provided by duplex sequencing, a partially double-stranded adapter having UMI (unique molecular identifiers) was ligated to double-stranded DNA.
[0286] To improve (and reduce costs) the robustness of duplex sequencing of EM-Seq against oxidation efficiency, conversion-resistant adapters can be used to enhance the consistency of sequencing library quality. The conversion-resistant adapters contain only unmodified bases and enable full base conversion of the adapters. An example of a conversion-resistant adapter is shown in Panel A of Figure 4.
[0287] In the case of no conversion, the sequencing library generated using these conversion-resistant adapters can be amplified and sequenced using a set of PCR and sequencing primers that match the original adapter sequence. In the case of conversion, the sequencing library can be amplified and sequenced using PCR and sequencing primers that match the converted adapter sequence as shown in Panel B of Figure 4.
[0288] One functional example set of a conversion resistance adapter, PCR primers, and sequencing primers was tested (Figure 6). The sequencing library yields of libraries generated using either a conversion resistance adapter or a 5mC-containing adapter are shown. Since the TET-mediated oxidation step was not performed, all Cs and 5mCs were susceptible to conversion from C to U. The 5mC-containing adapter system was more efficient without conversion, while the conversion resistance adapter required conversion so that the conversion resistance adapter system could be amplified using conversion-specific PCR primers. The DNA sequences for these conversion-specific PCR primers are listed in Table 1.
[0289]
Table 1-1
[0290]
Table 1-2
[0291] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided within the specification. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many changes, variations, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative ratios described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be utilized in practicing the present invention. Therefore, the present invention is intended to cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and methods and structures within the scope of these claims and their equivalents are intended to be embraced thereby.
Claims
1. A computer system for performing methylation sequencing on nucleic acid molecules of a biological sample, wherein the methylation sequencing comprises: a) ligating a nucleic acid adapter to the nucleic acid molecule, wherein the nucleic acid molecule comprises an unconverted nucleic acid, and the nucleic acid adapter comprises bases of guanine, thymine, adenine, and cytosine, and is a conversion-resistant adapter that does not contain 5-methylcytosine (5mC)-containing bases and 5-hydroxymethylcytosine (5hmC)-containing bases; b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid, wherein the enzymatic conversion comprises performing a series of ten-eleven translocation / apolipoprotein B mRNA editing enzyme (TET / APOBEC) enzymatic conversion or Tet-assisted pyridine borane sequencing (TAPS); c) amplifying at least partially the converted nucleic acid by polymerase chain reaction (PCR) to generate an amplified converted nucleic acid, wherein the PCR comprises using a sequencing primer corresponding to the sequence of the conversion-resistant adapter; d) contacting the amplified converted nucleic acid with a nucleic acid probe that is at least partially complementary to a panel of pre-identified CpG, CHG, or CHH loci to enrich for sequences corresponding to the panel of pre-identified CpG, CHG, or CHH loci and generate an enriched converted nucleic acid; e) determining the nucleic acid sequence of the enriched converted nucleic acid at a depth greater than 100x; f) comparing the nucleic acid sequence of the enriched converted nucleic acid with a reference nucleic acid sequence of a pre-identified panel of CpG, CHG, or CHH loci to determine the methylation profile of the nucleic acid molecules of the biological sample.
2. The computer system according to claim 1, wherein the nucleic acid molecule is plasma cell-free deoxyribonucleic acid (cfDNA).
3. The computer system according to claim 1, wherein the enzymatic conversion comprises Tet-assisted pyridine borane sequencing (TAPS) or chemically assisted pyridine borane sequencing (CAPS).
4. The computer system according to claim 1, wherein the nucleic acid adapter further comprises a unique dual index (UDI) sequence.
5. The computer system according to claim 4, wherein the length of the UDI sequence is 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, or 12 nucleotides.
6. The computer system according to claim 1, wherein the step of amplifying the converted nucleic acid comprises using a primer containing a unique dual index (UDI) sequence.
7. The computer system according to claim 1, wherein the nucleic acid probe is an unmethylated nucleic acid probe.
8. The computer system according to claim 1, wherein the nucleic acid probe hybridizes to a target region of interest that matches an unmethylated cytosine at a CpG site within a reference nucleic acid sequence.
9. The computer system according to claim 8, wherein the nucleic acid probe includes a target region of interest that matches a methylated cytosine at a CpG site within a reference nucleic acid sequence.
10. The computer system according to claim 1, wherein the nucleic acid probe is a mixture of chemically or enzymatically modified methylated nucleic acid probes or unmethylated nucleic acid probes.
11. The computer system according to claim 1, wherein one or more cytosines at the CG locus of the concentrated converted nucleic acid are converted to thymine, and all cytosines at the CH locus of the concentrated converted nucleic acid are converted to thymine.
12. A computer system for determining a targeted methylation pattern within a nucleic acid molecule of a biological sample derived from a subject, wherein the determination of the targeted methylation pattern comprises a) ligating a nucleic acid adapter to the nucleic acid molecule, wherein the nucleic acid molecule contains an unconverted nucleic acid, and the nucleic acid adapter is a conversion-resistant adapter that contains each of the bases guanine, thymine, adenine, and cytosine and does not contain a 5-methylcytosine (5mC)-containing base or a 5-hydroxymethylcytosine (5hmC)-containing base; b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid, wherein the enzymatic conversion comprises performing a series of ten-eleven translocation / apolipoprotein B mRNA editing enzyme (TET / APOBEC) enzymatic conversions or Tet-assisted pyridine borane sequencing (TAPS). c) amplifying at least partially the converted nucleic acid by polymerase chain reaction (PCR) to generate an amplified converted nucleic acid, wherein the PCR comprises the use of a sequencing primer corresponding to the sequence of the conversion resistance adapter, d) contacting the amplified converted nucleic acid with a nucleic acid probe that is at least partially complementary to a panel of pre-identified CpG, CHG, or CHH loci to enrich for sequences corresponding to the panel of pre-identified CpG, CHG, or CHH loci and generate an enriched converted nucleic acid, e) determining the nucleic acid sequence of the enriched converted nucleic acid at a depth greater than 100x, f) comparing the nucleic acid sequence of the enriched converted nucleic acid with a reference nucleic acid sequence of a pre-identified panel of CpG, CHG, or CHH loci to determine the methylation pattern of the nucleic acid molecules of the biological sample from the subject, a computer system comprising the steps.
13. The computer system according to claim 12, wherein the step of determining the nucleic acid sequence of the enriched converted nucleic acid comprises duplex-like error correction.
14. The computer system according to claim 12, wherein the pre-identified panel of CpG, CHG, or CHH loci comprises loci associated with transcription factor start sites.
15. The computer system according to claim 12, wherein the targeted methylation pattern comprises hemimethylated CpG loci.
16. A computer system for determining the methylation profile of a cell-free deoxyribonucleic acid (cfDNA) sample from a subject, wherein the determination of the methylation profile comprises a) ligating a nucleic acid adapter to the cfDNA, wherein the cfDNA contains unconverted nucleic acid and the nucleic acid adapter is a conversion-resistant adapter that contains each of the bases guanine, thymine, adenine, and cytosine and does not contain 5-methylcytosine (5mC)-containing bases and 5-hydroxymethylcytosine (5hmC)-containing bases, b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid, wherein the enzymatic conversion comprises performing a series of ten-eleven translocation / apolipoprotein B mRNA editing enzyme (TET / APOBEC) enzymatic conversions or Tet-assisted pyridine borane sequencing (TAPS). c) amplifying at least partially the converted nucleic acid by polymerase chain reaction (PCR) to generate an amplified converted nucleic acid, wherein the PCR includes the use of a sequencing primer corresponding to the sequence of the conversion resistance adapter, d) contacting the amplified converted nucleic acid with a nucleic acid probe that is at least partially complementary to a panel of pre-identified CpG, CHG, or CHH loci to enrich for sequences corresponding to the panel of pre-identified CpG, CHG, or CHH loci and generate an enriched converted nucleic acid, e) determining the nucleic acid sequence of the enriched converted nucleic acid at a depth greater than 100x, f) comparing the nucleic acid sequence of the enriched converted nucleic acid with a reference nucleic acid sequence of a pre-identified panel of CpG, CHG, or CHH loci to determine a methylation profile of a cfDNA sample derived from a subject, a computer system comprising.
17. The computer system according to claim 16, wherein the determination of the methylation profile further includes a step of identifying a cfDNA sample derived from a tissue, a step of identifying somatic mutations in the cfDNA sample, a step of estimating the nucleosome position of the cfDNA sample, a step of identifying a variable methylation region of the cfDNA sample, or a step of identifying a haplotype block of the cfDNA sample.
18. A computer system for generating a classifier, wherein the generation of the classifier comprises: a) ligating a nucleic acid adapter to each nucleic acid molecule of a biological sample from a healthy subject and a biological sample from a cancer subject, wherein the nucleic acid adapter is a conversion resistance adapter that contains each base of guanine, thymine, adenine, and cytosine and does not contain a 5-methylcytosine (5mC)-containing base and a 5-hydroxymethylcytosine (5hmC)-containing base, b) enzymatically converting unmethylated cytosine to uracil within the nucleic acid molecule to generate a converted nucleic acid, wherein the enzymatic conversion includes performing a series of ten-eleven translocation / apolipoprotein B mRNA editing enzyme (TET / APOBEC) enzymatic conversions or Tet-assisted pyridine borane sequencing (TAPS). c) amplifying at least partially the converted nucleic acid by polymerase chain reaction (PCR) to generate an amplified converted nucleic acid, wherein the PCR involves using a sequencing primer corresponding to the sequence of the conversion resistance adapter; d) contacting the amplified converted nucleic acid with a nucleic acid probe that is at least partially complementary to a panel of pre-identified CpG, CHG, or CHH loci to enrich for sequences corresponding to the panel of pre-identified CpG, CHG, or CHH loci and generate an enriched converted nucleic acid; e) determining the nucleic acid sequence of the enriched converted nucleic acid at a depth greater than 100x; f) comparing the nucleic acid sequence of the enriched converted nucleic acid with a reference nucleic acid sequence of a pre-identified panel of CpG, CHG, or CHH loci to obtain a set of measurements of input features representing the methylation profiles of healthy and cancer subjects; g) training a machine learning model to generate a classifier that discriminates between healthy and cancer subjects, the classifier being a computer system. **Claim 19** The computer system according to claim 18, wherein the pre-identified panel of CpG, CHG, or CHH loci includes loci associated with transcription start sites. **Claim 20** The computer system according to claim 18, wherein the generation of the classifier further includes determining loci of hemimethylated CpG, CHG, or CHH. **Claim 21** The computer system according to claim 18, wherein the generation of the classifier further includes identifying the tissue origin of the nucleic acid molecule. **Claim 22** The computer system according to claim 18, wherein the generation of the classifier further includes identifying the genomic location and fragment length of the nucleic acid molecule. **Claim 23** The input features are selected from the methylation percentage of base units for CpG, the methylation percentage of base units for CHG, the methylation percentage of base units for CHH, the number or ratio of fragments with different numbers or ratios of methylated CpGs in a region, conversion efficiency, hypomethylated blocks, methylation levels of CpG, methylation levels of CHH, methylation levels of CHG, fragment length, fragment midpoint, methylation levels of chrM, methylation levels of LINE1, methylation levels of ALU, dinucleotide coverage, uniformity of coverage, overall average CpG coverage, and average coverage in CpG islands, CGI shelves, and CGI shores, of the computer system according to claim 18.
24. A tumor detection kit comprising a reagent and instructions for detecting a tumor signal for implementing the computer system according to any one of claims 1 to 23.
25. The tumor detection kit according to claim 24, wherein the reagent is selected from the group consisting of a primer set, PCR reaction components, sequencing reagents, enzyme conversion reagents, and library preparation reagents.
Citation Information
Patent Citations
Methods and systems for determining cancer status
JP2018508228A
Compositions and Methods for Analyzing Modified Nucleotides
US20170198344A1
Methods for targeted nucleic acid sequence enrichment with applications to error corrected nucleic acid sequencing
WO2018175997A1
Methods and systems for evaluating DNA methylation in cell-free DNA
WO2019006269A1