Methods and systems for high-depth sequencing of methylated nucleic acid
The method addresses DNA degradation and high cost issues in bisulfite-based methylation sequencing by using enzymatic conversion and duplex sequencing, enhancing accuracy and sensitivity for cfDNA methylation analysis.
Patent Information
- Application Number
- JP2025122620
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-05
AI Technical Summary
Existing methylation sequencing methods, particularly bisulfite-based approaches, cause significant DNA degradation and require excessive sample input, leading to limited sensitivity and high costs for analyzing cell-free DNA (cfDNA) methylation patterns, which are crucial for early cancer detection and diagnosis.
A method involving minimally disruptive enzymatic conversion of unmethylated cytosines to uracils, followed by PCR amplification, probe enrichment for specific CpG or CH loci, and deep sequencing to determine methylation profiles, utilizing unique molecular identifiers and duplex sequencing for error correction.
This approach preserves cfDNA integrity, enhances sequencing accuracy, reduces sample requirements, and improves the sensitivity and cost-effectiveness of methylation analysis, enabling more reliable disease detection and classification.
Smart Images

Figure 2025165986000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 62 / 855,795, filed May 31, 2019, which is incorporated herein by reference in its entirety.
[0002] Citation by reference All publications, patents, and patent applications mentioned in this specification, including the examples, are herein incorporated by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions therein, will control. [Background technology]
[0003] Due to the stability and role of DNA in normal differentiation and diseases such as cancer, DNA methylation can represent tumor characteristics and phenotypic states, making it potentially useful in personalized medicine. Abnormal DNA methylation patterns occur early in cancer development, facilitating early cancer detection. In fact, aberrant DNA methylation is one of the hallmarks of cancer and has been linked to all aspects of cancer, from tumor initiation to cancer progression and metastasis. These characteristics have led to the use of DNA methylation patterns in many recent approaches to diagnose cancer. Specifically, cell-free DNA (cfDNA) is fragmented DNA present in the circulation, and its fragmentation pattern is a useful and informative biological signal. In contrast, genomic DNA is artificially fragmented in vitro for use in library preparation, making the fragmentation pattern of genomic DNA less relevant in diagnostic methods.
[0004] DNA methylation is a covalent modification of DNA and a stable genetic trait that can play important roles in repressing gene expression and regulating chromatin structure. In humans, DNA methylation occurs primarily at cytosine residues in CpG dinucleotides. Unlike other dinucleotides, CpGs are not evenly distributed throughout the genome and can be concentrated in DNA regions rich in short CpGs, known as CpG islands. Typically, the majority of CpG sites in the genome, approximately 70–75%, are methylated. However, methylation patterns vary across cell types, reflecting their role in regulating cell-type-specific gene expression. In this way, the cellular methylome can program the cell's terminal differentiation state, for example, to become neurons, muscle cells, or immune cells.
[0005] Furthermore, various cell subtypes within a tissue can exhibit distinct methylation patterns. In cancer cells, CpG methylation can be deregulated, and aberrant methylation patterns are some of the earliest events in tumorigenesis. The methylation profile of a given cancer type is most similar to that of the tissue from which it originates. Therefore, aberrant methylation signatures on cfDNA fragments can be used to differentiate cancer cells from normal cells and identify their tissue type of origin. While global CpG methylation levels are typically reduced in cancer cells, the average methylation level (or % methylation) at specific loci may vary at specific CpG sites in cancer cells compared to paired normal cells. By profiling differentially methylated CpGs (DMCs, single sites) or regions (DMRs, multiple sites within a local region) between normal and diseased cells, disease biomarkers can be identified. This approach led to the development of the SEPT9 gene methylation assay (Epi proColon). This assay is the first FDA-approved blood-based diagnostic for colorectal cancer (CRC).
[0006] Bisulfite conversion, or bisulfite sequencing, has become a widely used method for DNA methylation analysis. Bisulfite sequencing is a convenient and effective method for mapping DNA methylation to individual bases. Unfortunately, bisulfite conversion is a harsh and destructive process for cfDNA, degrading more than 90% of the sample DNA. The two main approaches to constructing bisulfite sequencing libraries are (1) bisulfite conversion of DNA before library construction, which requires the construction of a single-stranded DNA library, and (2) bisulfite conversion of DNA after double-stranded adapter ligation. Both involve severe DNA degradation, which can be particularly problematic for cfDNA, which is present in very low concentrations in plasma and is a limiting resource for liquid biopsy applications. In ssDNA libraries, some degraded cfDNA can be retained within the library, but endpoint information on the degraded fragments is lost. Such libraries limit the ability to use cfDNA endpoint or fragment length information to test DNA methylation. In dsDNA libraries, bisulfite-cleaved cfDNA inserts are lost from the library, but endpoint information on viable cfDNA inserts is retained. This requires excessive blood draws to achieve deep, specific coverage of the genome, or limits analytical performance to only low, specific coverage.
[0007] The advent of next-generation DNA sequencing has advanced clinical medicine and basic research. However, while this technique can generate hundreds of billions of nucleotides of DNA sequence in a single experiment, it has an error rate of approximately 1%, resulting in hundreds of millions of sequencing errors. While such errors are acceptable for some applications, they pose a significant problem for "deep sequencing" of genetically heterogeneous mixtures such as tumors or mixed microbial populations.
[0008] Existing methods require two different sequencing assays and two different cfDNA pools to analyze cfDNA variants and cfDNA methylation status. These methods can be prohibitively expensive in terms of plasma / cfDNA input volume and associated costs. Additionally, bisulfite-mediated DNA destruction can reduce the sensitivity of variant-calling methods that operate on bisulfite-converted DNA sequencing data (compared to enzymatic conversion). Therefore, improved methods for analyzing cfDNA methylation are needed to maintain the integrity of sample nucleic acids and improve the accuracy of methylation status analysis at the whole-genome or targeted level. Summary of the Invention
[0009] The methods and systems provided herein address the limitations of bisulfite-based methylation sequencing by improving the quality and accuracy of nucleic acid methylation sequencing, which is used for disease detection. Increased accuracy and complete information about methylation status can lead to higher quality feature generation used in machine learning models and classifier generation.
[0010] In a first aspect, there is provided a method for performing methylation sequencing of a nucleic acid sample, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to the nucleic acid molecule, wherein the nucleic acid molecule comprises unconverted nucleic acid; b) converting unmethylated cytosines to uracils in said nucleic acid molecule using a minimally disruptive conversion method, thereby producing a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction, thereby producing an amplified converted nucleic acid; d) probing the amplified converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby producing probed converted nucleic acid; e) determining the nucleic acid sequence of the probed converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequences of the probed converted nucleic acids with reference nucleic acid sequences of a pre-identified panel of CpG or CH loci to determine the methylation profile of the nucleic acid molecules of the biological sample.
[0011] In one embodiment, the nucleic acid molecule is plasma cfDNA.
[0012] In one embodiment, the minimally disruptive conversion method comprises enzymatic conversion, TAPS, or CAPS.
[0013] In one embodiment, the unique molecular identifier is 4-6 bp in length and has a 5' thymidine overhang.
[0014] In one embodiment, the nucleic acid adapter further comprises a unique dual index (UDI) sequence, hi one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0015] In one embodiment, amplifying the converted nucleic acid comprises using a primer comprising a unique dual index (UDI) sequence, hi one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0016] In one embodiment, the nucleic acid adapter is a conversion resistant adapter that includes guanine, thymine, adenine, and cytosine bases but does not include 5mC- or 5hmC-containing bases.
[0017] In one embodiment, the nucleic acid probe is an unmethylated nucleic acid probe.
[0018] In one embodiment, the nucleic acid probe hybridizes to a target region of interest that matches an unmethylated cytosine at a CpG site in a reference nucleic acid sequence.
[0019] In one embodiment, the nucleic acid probe comprises a target region of interest that matches a methylated cytosine at a CpG site in a reference nucleic acid sequence.
[0020] In one embodiment, the nucleic acid probes are a mixture of chemically or enzymatically modified methylated or unmethylated nucleic acid probes.
[0021] In one embodiment, one or more cytosines in a CG context of said probed converted nucleic acid are converted to thymine, and all cytosines in a CH context of said probed converted nucleic acid are converted to thymine.
[0022] In one embodiment, the conversion of unmethylated cytosine to uracil comprises a series of TET / APOBEC enzyme conversions.
[0023] In one embodiment, the conversion of unmethylated cytosine to uracil comprises TAPS.
[0024] In a second aspect, there is provided a method for determining targeted methylation patterns in nucleic acid molecules of a biological sample from a subject, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to cfDNA, wherein the cfDNA comprises unconverted nucleic acid; b) enzymatically converting unmethylated cytosines to uracils within the nucleic acid molecule to produce converted nucleic acids; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to said pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid to reference nucleic acid sequences of the pre-identified panel of CpG or CH loci to determine a methylation profile of a cell-free DNA (cfDNA) sample from the subject.
[0025] In one embodiment, determining the nucleic acid sequence of the converted nucleic acid comprises duplex-like error correction.
[0026] In one embodiment, the nucleic acid adapter is a conversion resistant adapter that includes guanine, thymine, adenine, and cytosine bases but does not include 5mC- or 5hmC-containing bases.
[0027] In one embodiment, said pre-identified panel of CpG or CH loci comprises loci associated with transcription factor start sites.
[0028] In one embodiment, said targeted methylation pattern comprises hemimethylated CpG loci.
[0029] In a third aspect, there is provided a method for determining a methylation profile of a cell-free DNA (cfDNA) sample from a subject, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to cfDNA, wherein the cfDNA comprises unconverted nucleic acid; b) enzymatically converting unmethylated cytosines to uracils within the nucleic acid molecule to produce converted nucleic acids; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to said pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid to reference nucleic acid sequences of the pre-identified panel of CpG or CH loci to determine a methylation profile of a cell-free DNA (cfDNA) sample from the subject.
[0030] In one embodiment, the nucleic acid adapter is a conversion resistant adapter that includes guanine, thymine, adenine, and cytosine bases but does not include 5mC- or 5hmC-containing bases.
[0031] In one embodiment, the unique molecular identifier is 4-6 bp in length and has a 5' thymidine overhang.
[0032] In one embodiment, the nucleic acid adapter further comprises a unique dual index (UDI) sequence, hi one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0033] In one embodiment, the method further comprises identifying a cfDNA sample from the tissue, identifying somatic mutations in the cfDNA sample, estimating nucleosome positions in the cfDNA sample, identifying variably methylated regions in the cfDNA sample, or identifying haplotype blocks in the cfDNA sample.
[0034] Further provided herein is a method for methylation sequencing using duplex sequencing. Duplex sequencing is a tag-based error correction method that can improve sequencing accuracy, e.g., methylation sequencing accuracy. In this method, adapters are ligated onto nucleic acid templates and amplified using PCR. In one embodiment, the adapters include a primer sequence and a random 12-bp index. Deep sequencing provides consensus sequence information from all unique molecular tags. Based on the molecular tags and sequencing primers, the duplex sequences can be aligned to determine the true sequence of the DNA. Advantages of duplex sequencing include a very low error rate and the ability to detect and remove PCR amplification errors. Duplex sequencing does not require an additional library preparation step after the addition of adapters.
[0035] In one embodiment, a method for performing methylation sequencing on nucleic acid molecules of a biological sample comprises: a) preparing a methylation sequencing library from cfDNA fragments of nucleic acid molecules, i) ligating double-stranded adapters to cfDNA fragments; ii) ligating a dual unique molecular identifier to the cfDNA fragment; and iii) preparing a methylation sequencing library from the nucleic acid molecule cfDNA by converting unmethylated cytosines in the cfDNA fragments to uracils using a minimally disruptive conversion method; b) enriching the methylation sequencing library for sequences corresponding to CpG or CH loci, thereby generating an enriched methylation sequencing library; c) sequencing the enriched methylation sequencing library at a depth of greater than 100x using single-end or paired-end reads, thereby generating single-end or paired-end read sequencing fragments; d) for each sequencing fragment of the paired-end reads, correcting sequencing errors within the overlapping region of the paired-end reads; e) collapsing the sequencing fragments into concatenated read families to correct errors resulting from PCR and sequencing; f) collapsing the concatenated read families into dual read families to identify methylation discrepancies relative to the predicted methylation status of symmetric CpG loci in the nucleic acid molecule.
[0036] In one embodiment, the minimally disruptive conversion comprises an enzymatic conversion, TAPS, or CAPS.
[0037] In a fourth aspect, there is provided a method for generating a classifier, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to each nucleic acid molecule of a biological sample from a healthy subject and a biological sample from a cancer subject; b) converting unmethylated cytosines to uracils in said nucleic acid molecule using a minimally disruptive conversion method, thereby producing a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction, thereby producing an amplified converted nucleic acid; d) probing the amplified converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby producing probed converted nucleic acid; e) determining the nucleic acid sequence of the probed converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequences of said probed converted nucleic acids with reference nucleic acid sequences of said pre-identified panel of CpG or CH loci to obtain a set of measured input features representative of the methylation profiles of healthy and cancer subjects; g) training a machine learning model to generate a classifier that distinguishes between healthy and cancerous subjects.
[0038] In one embodiment, said pre-identified panel of CpG or CH loci comprises loci associated with transcription start sites.
[0039] In one embodiment, the method further comprises determining hemimethylated CpG or CH loci.
[0040] In one embodiment, the method further comprises identifying the tissue origin of the nucleic acid molecule.
[0041] In one embodiment, the method further comprises identifying the genomic location and fragment length of the nucleic acid molecule.
[0042] In one embodiment, the unique molecular identifier is 4-6 bp in length and has a 5' thymidine overhang.
[0043] In one embodiment, the nucleic acid adapter further comprises a unique dual index (UDI) sequence, hi one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0044] In one embodiment, amplifying the converted nucleic acid comprises using a primer comprising a unique dual index (UDI) sequence, hi one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0045] In one embodiment, the input features are selected from base-wise methylation % for CpG, base-wise methylation % for CHG, base-wise methylation % for CHH, number or percentage of observed fragments with different number or percentage of methylated CpGs in a region, conversion efficiency (e.g., 100 average methylation % for CHH), hypomethylated blocks, CPG methylation level, CHH methylation level, CHG methylation level, fragment length, fragment midpoint, chrM methylation level, LINE1 methylation level, ALU methylation level, dinucleotide coverage (e.g., normalized dinucleotide coverage), coverage uniformity (e.g., unique CPG sites at 1x and 10x average genome coverage (e.g., for s4 runs)), overall average CpG coverage (e.g., depth), and average coverage at CpG islands, CGI shelves, and CGI shores.
[0046] In a fifth aspect, there is provided a classifier for distinguishing between healthy and cancerous individuals, the classifier comprising a set of measurements representing methylation profiles from methylation sequencing data of each of healthy and cancerous subjects, the measurements being used to generate a set of features corresponding to characteristics of the methylation profiles, the set of features being input into a machine learning or statistical model, wherein the machine learning or statistical model provides a feature vector useful as a classifier for distinguishing between populations of healthy and cancerous individuals.
[0047] In a sixth aspect, a method for detecting cancer in a population of subjects is provided, the method comprising: a) assaying nucleic acids of a biological sample from a subject using minimally destructive targeted conversion methyl sequencing to obtain a methylation profile of the nucleic acid; b) classifying the biological samples by inputting the methylation profile into a training algorithm that classifies each sample between healthy and cancerous subjects; and c) if the training algorithm classifies the biological sample as negative for cancer with a particular confidence value, outputting a report on a computer screen identifying the biological sample as negative for cancer.
[0048] In one example, the cancer is colon cancer.
[0049] In a seventh aspect, the present disclosure provides a system for classifying individuals based on methylation status, the system comprising: a) a computer readable medium product comprising a classifier, the classifier comprising a set of measurements representing methylation profiles from methylation sequencing data of each of healthy and cancer subjects, the measurements being used to generate a set of features corresponding to characteristics of each of the methylation profiles of the healthy and cancer subjects, the set of features being input into a machine learning or statistical model, the machine learning or statistical model providing a feature vector useful as a classifier for distinguishing between populations of healthy and cancer individuals; b) one or more processors for executing the instructions stored on said computer-readable media product.
[0050] In one example, the system includes a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first-order polynomial kernel support vector machine classifier, a second-order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimum problem optimization algorithm classifier, a naive Bayes algorithm classifier, and a non-negative matrix factorization (NMF) prediction algorithm classifier.
[0051] In one embodiment, the system comprises means for performing any of the above methods.
[0052] In one embodiment, the system comprises one or more processors configured to perform any of the above methods.
[0053] In one embodiment, the system comprises modules for performing each of the steps of any of the above methods.
[0054] In another aspect, the present disclosure provides a method for monitoring minimal residual disease status of a subject previously treated for a disease, the method comprising determining a methylation profile as described herein as a baseline methylation status and repeating the analysis to determine said methylation profile at one or more predetermined time points, wherein a change from baseline indicates a change in the minimal residual disease status at baseline in the subject.
[0055] In another aspect, the present disclosure provides a method for monitoring minimal residual disease status of a subject previously treated for a disease, the method comprising: a) determining a baseline methylation profile of a biological sample obtained from a subject at baseline methylation status; b) determining a test methylation profile of a biological sample obtained from the subject at one or more predetermined time points after the baseline methylation state; and c) determining a change in the test methylation profile as compared to the baseline methylation profile, said change indicating a change in the minimal residual disease status of the subject.
[0056] In some embodiments, the disease is cancer, hi some embodiments, the disease is colon cancer.
[0057] In another aspect, the present disclosure provides a method for monitoring minimal residual disease status of a subject previously treated for colorectal cancer, the method comprising detecting methylated fragments in a biological sample from the subject, wherein the methylated fragments in the biological sample indicate a change in the subject's baseline minimal residual disease status for colorectal cancer.
[0058] In some embodiments, the minimal residual disease status is selected from response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.
[0059] In another aspect, a method for determining a subject's response to treatment is provided.
[0060] In another aspect, a method for monitoring tumor burden in a subject is provided.
[0061] In another aspect, a method is provided for detecting residual tumor after surgery in a subject.
[0062] In another aspect, a method for detecting recurrence in a subject is provided.
[0063] In another aspect, a method is provided for use as a secondary screen for subjects.
[0064] In another aspect, a method is provided for use as a primary screen for subjects.
[0065] In another aspect, a method for monitoring cancer progression in a subject is provided.
[0066] In another aspect, the present disclosure provides a kit for tumor detection, comprising reagents for carrying out the aforementioned method and instructions for detecting tumor signals, e.g., methylation signals, such as primer sets, PCR reaction components, sequencing reagents, minimally disruptive conversion reagents, and library preparation reagents. [Brief explanation of the drawings]
[0067] [Figure 1] 1 is a flow diagram comparing conventional methyl bisulfite conversion and degradation with a modified method that preserves fragment length information as described herein. [Figure 2] FIG. 1 is a schematic diagram showing staggered adapters useful in the methods described herein. [Figure 3] FIG. 1 is a schematic diagram illustrating the duplex-like error correction method described herein. [Figure 4] PANEL A is a schematic diagram showing an example of a conversion resistant adapter, and PANEL B is the corresponding sequencing primer that matches the conversion adapter sequence to perfectly base-pair with a compatible PCR primer. [Figure 5] 1 illustrates a computer system that is programmed or otherwise configured to perform the methods provided herein. [Figure 6] 1 is a graph showing sequencing library yields for exemplary conversion-resistant adapter / primer systems, where higher sequencing yields due to conversion are observed for conversion-resistant adapters compared to 5mC-containing adapters. DETAILED DESCRIPTION OF THE INVENTION
[0068] Provided herein are methods that enable improved library preparation and sequencing of methylated regions for cfDNA methylation profiling. The methods address limitations of traditional methylation sequencing and profiling of nucleic acids in biological samples by improving coverage, coverage uniformity, resolution, and methylation data accuracy to support practical applications. Sequencing data obtained from the methods provided herein are useful for practical applications that use methylation profiling data to classify or stratify individual populations. Such population classification and stratification can include identifying individuals with disease, staging disease progression, or response to specific disease treatments.
[0069] I. Definition As used herein, singular terms such as "a," "an," and "the" are intended to include both the singular and the plural unless the context clearly indicates otherwise.
[0070] The terms "plasma cell-free DNA," "circulating free DNA," "cell-free DNA," or "cfDNA" can refer to DNA molecules circulating in the cell-free portion of blood. Circulating nucleic acids in blood originate from necrotic or apoptotic cells, and in diseases such as cancer, extremely high levels of nucleic acids are observed from apoptosis. In cancer, circulating DNA carries hallmarks of disease, including cancer gene mutations and microsatellite alterations. This circulating DNA is sometimes referred to as circulating tumor DNA (ctDNA). Viral genomic sequences, DNA, or RNA in plasma are potential biomarkers of disease.
[0071] In some embodiments, the cell-free fraction of blood is preferably serum or plasma. As used herein, the term "cell-free fraction" of a biological sample refers to a fraction of the biological sample that is substantially free of cells. As used herein, the term "substantially cell-free" may refer to a preparation derived from a biological sample that contains less than about 20,000 cells / ml, less than about 2,000 cells / ml, less than about 200 cells / ml, or less than about 20 cells / ml. Genomic DNA (gDNA) refers to the unfragmented DNA released from white blood cells that contaminates the cell-free fraction of blood. To mitigate gDNA-contaminated samples, a highly controlled sample processing workflow may be implemented, and specimens may be screened for the presence of gDNA.
[0072] As used herein, the term "diagnosing" or "diagnosis" in reference to a condition or outcome includes predicting or diagnosing a condition or outcome, determining a predisposition to a condition or outcome, monitoring a patient's treatment, diagnosing a patient's response to treatment, prognosis, progression, and response to a particular treatment of a condition or outcome.
[0073] As used herein, the term "position" refers to an identified strand nucleotide position within a nucleic acid molecule.
[0074] As used herein, the term "nucleic acid" refers to DNA, RNA, DNA / RNA chimeras or hybrids, which may be single-stranded (ss) or double-stranded (ds). Nucleic acids may be genomic, derived from the genome of a eukaryotic or prokaryotic cell, synthesized, cloned, amplified, or reverse transcribed. In certain embodiments of the methods and compositions, nucleic acid preferably refers to genomic DNA, as required by context.
[0075] As used herein, unless otherwise specified, the term "modified cytosine" refers to 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), formyl-modified cytosine, carboxy-modified cytosine, 5-carboxylcytosine (5caC), or a cytosine modified with any other chemical group.
[0076] As used herein, the terms "methylcytosine dioxygenase," "dioxygenase," or "oxygenase" refer to an enzyme that converts 5mC to 5hmC. Non-limiting examples of methylcytosine dioxygenases include TET1, TET2, TET3, and Naegleria TET. TET2 is an example of a methylcytosine dioxygenase that oxidizes at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, or at least 99% of all 5mC.
[0077] As used herein, the terms "conversion-resistant adapter" and "conversion-resistant primer" refer to nucleic acid molecules used as adapters or primers, respectively. Instead of incorporating modified nucleotide bases to prevent base conversion, conversion-resistant adapters or conversion-resistant primers incorporate only unmodified bases, thereby allowing full base conversion during the conversion reaction for methylation sequencing. "Unmodified bases" in the adapter / primer DNA sequence refer to the conventional guanine, adenine, cytosine, and thymine.
[0078] As used herein, the term "cytidine deaminase" refers to an enzyme that deaminates cytosine (C) to form uracil (U). Non-limiting examples of cytidine deaminases include the APOBEC family of cytidine deaminases, such as APOBEC3A. In any embodiment, a cytidine deaminase described herein may have a sequence that is at least 90% identical (e.g., at least 95% identical) to the amino acid sequence of GenBank Accession No. AKE33285.1, which is the sequence of human APOBEC3A. In some embodiments, a cytidine deaminase described herein converts unmodified cytosine to uracil with an efficiency of at least 95%, 98%, or 99%, preferably at least 99%.
[0079] As used herein, the term "glucosyltransferase" or "GT" refers to an enzyme that catalyzes the transfer of a β-D-glucosyl or α-D-glucosyl residue from UDP-glucose to a 5hmC residue to form 5ghmC. APOBEC can convert 5hmC to U at a lower rate than it converts C or 5mC to U. An example of a GT is T4-beta GT (βGT). In one example, a GT can be used in conjunction with a dioxygenase. This combination ensures that 5hmC deamination is blocked, such that less than 5%, less than 3%, or less than 1% of 5hmC is converted to U by the deaminase. In another example, a GT can be used with a dioxygenase in the same reaction mixture with DNA, where the dioxygenase converts 5mC to 5hmC and 5caC, and the GT converts any remaining 5hmC to 5ghmC, ensuring that only cytosines are deaminated.
[0080] As used herein, the terms "portion" and "aliquot" with respect to a nucleic acid sample are intended to have the same meaning and can be used interchangeably.
[0081] As used herein, the term "comparing" refers to analyzing two or more sequences relative to one another. Optionally, comparison may be performed by aligning two or more sequences relative to one another, so that nucleotides located at corresponding positions are aligned relative to one another.
[0082] As used herein, the term "reference sequence" refers to the sequence of the fragment being analyzed. The reference sequence can be obtained from a public database or can be sequenced separately as part of an experiment. In some cases, the reference sequence can be hypothetical, whereby the reference sequence can be computationally deaminated (i.e., Cs are changed to Us or Ts, etc.) to allow for sequence comparison to be performed.
[0083] As used herein, the terms "G," "A," "T," "U," "C," "5mC," "5fC," "c5aC," "5hmC," and "5ghmC" refer to guanidine (G), adenine (A), thymine (T), uracil (U), cytosine (C), 5-methylcytosine, 5-formylcytosine, 5-carboxylcytosine (5caC), 5-hydroxymethylcytosine, and 5-glucosylhydroxymethylcytosine, respectively. For clarity, C, 5fc, 5caC, 5mc, and 5ghmC are each different moieties.
[0084] The term "minimal residual disease" or "MRD" refers to the small number of cancer cells remaining in the body after cancer treatment. MRD testing can be performed to determine whether cancer treatment has worked and to plan further treatment. Various metrics can be used to assess MRD, including, but not limited to, response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.
[0085] The term "next generation sequencing" or "NGS" generally applies to sequencing libraries of genomic fragments that are less than 1 kb in size.
[0086] As used herein, the term "healthy" refers to a disease-free subject or a sample derived therefrom. While health is a dynamic state, the term can also refer to the pathological state of a subject lacking a referenced disease state, e.g., cancer. In one example, when referring to a methylation profile for classifying a cancer subject, the term "healthy" refers to an individual lacking cancer, such as CRC. While other diseases or health conditions may be present in this subject, the term "healthy" may refer to the absence of a manifested disease in order to compare or classify subjects with and without disease states, and samples derived therefrom.
[0087] As used herein, the term "threshold" generally refers to a value selected to discriminate, separate, or differentiate between two subject populations. In some embodiments, the threshold discriminates methylation status between diseased (e.g., malignant) and non-diseased (e.g., healthy) states. In some embodiments, the threshold discriminates disease stage (e.g., stage 1, stage 2, stage 3, or stage 4). The threshold can be set according to the disease of interest, for example, based on an initial analysis of a training set, or can be computationally determined for a set of inputs with known characteristics (e.g., healthy, disease, or disease stage). Thresholds can also be set for gene regions according to the predictive value of methylation at specific sites. Thresholds can be different for different methylation sites, and data from multiple sites can be combined in the final analysis.
[0088] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the teachings of the present invention, some exemplary methods and materials are described herein.
[0089] The citation of any publication is for its disclosure prior to the filing date of the present application and should not be construed as an admission that the present claims are not entitled to antedate such publication by virtue of prior invention. Further, the publication dates provided may be different from the actual publication dates that can be independently confirmed.
[0090] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features that can be readily separated from or combined with the features of any of the other various embodiments without departing from the scope or spirit of the present teachings. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.
[0091] All patents and publications referred to herein, including all sequences disclosed within those patents and publications, are expressly incorporated by reference.
[0092] II. Targeted methylation sequencing In the method of targeted methylation sequencing, the target region in biological sample such as cfDNA is analyzed to determine the methylation status of target gene sequence.In some embodiments, this target region comprises the consecutive nucleotides of the target region of interest, such as at least about 16 consecutive nucleotides of the target region of interest, or hybridizes to this consecutive nucleotide under stringent conditions.In various examples, targeted sequencing can be achieved by using hybridization capture and amplicon sequencing methods.
[0093] A. Hybridization Capture The hybridization method provided herein can be used for various formats of nucleic acid hybridization, such as in-solution hybridization, hybridization on a solid support (e.g., Northern hybridization, Southern hybridization, and in situ hybridization on membranes, microarrays, and cell / tissue slides). This method is particularly suitable for in-solution hybrid capture for target enrichment of certain genomic DNA sequences (e.g., exons) used in targeted next-generation sequencing. In hybrid capture techniques, cell-free nucleic acid samples undergo library preparation. As used herein, "library preparation" includes end repair, A-tailing, adapter ligation, or other preparations performed on cell-free DNA to enable subsequent DNA sequencing. In some examples, the prepared cell-free nucleic acid library sequences include adapters, sequence tags, and index barcodes ligated to the cell-free nucleic acid sample molecules. Various commercially available kits are also available to facilitate library preparation for next-generation sequencing techniques. Next-generation sequencing library construction may involve preparing nucleic acid targets using a series of tailored enzymatic reactions to generate a collection of DNA fragments of specific sizes for high-throughput sequencing. Advances and developments in various library preparation techniques have expanded the application of next-generation sequencing to fields such as transcriptomics and epigenetics.
[0094] As a result of improvements in sequencing technology, changes and improvements have been observed in library preparation. As used herein, next generation sequencing library preparation kits include kits developed by companies such as Agilent, Bioo Scientific, Kapa Biosystems, New England Biolabs, Illumina, Life Technologies, Pacific Biosciences, and Roche.
[0095] In various examples for targeted capture gene panels, various library preparation kits can be selected from Nextera Flex (Illumina), IonAmpliseq (Thermo Fisher Scientific), Genexus (Thermo Fisher Scientific), Agilent ClearSeq (Illumina), Agilent SureSelect Capture (Illumina), Archer FusionPlex (Illumina), BioScientific NEXTflex (Illumina), IDT xGen (Illumina), Illumina TruSight (Illumina), Nimblegene SeqCap (Illumina), and Qiagen GeneRead (Illumina).
[0096] In some embodiments, the hybrid capture method is performed on the prepared library sequence using a specific probe. As used herein, the term "specific probe" may refer to a probe specific to a known methylation site. In some embodiments, the specific probe is designed based on the use of the human genome as a reference sequence and a specific genomic region known to have a methylation site as a target sequence. Specifically, the genomic region known to have a methylation site may include at least one of the following: a promoter region, a CpG island region, a CGI shore region, and an imprinted gene region. Therefore, when hybrid capture is performed using the specific probe of some embodiments, sequences in the sample genome that are complementary to the target sequence, for example, regions in the sample genome known to have a methylation site (also referred to herein as "specific genomic regions"), can be efficiently captured.
[0097] In one example, the methylation regions described herein are used to design specific probes. In some embodiments, the specific probes are designed using commercially available methods, such as the eArray system. The length of the probe can be long enough to hybridize with sufficient specificity to the desired methylation region. In various examples, the probe is a 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, or 20-mer.
[0098] The target region for methylation analysis can be screened by utilizing database resources (such as gene ontology). According to the principle of complementary base pairing, a single-stranded capture probe can be complementarily combined with a single-stranded target sequence to successfully capture the target region. In some embodiments, the designed probe can be designed as a solid capture chip (the probe is immobilized on a solid support) or a liquid capture chip (the probe is free in a liquid). However, due to limiting factors such as probe length, probe density, and high cost, solid capture chips are rarely used, while liquid capture chips are more frequently used.
[0099] In some embodiments, GC-rich sequences (with an average GC base content greater than 60%) may result in reduced capture efficiency due to the molecular structure of C and G bases compared to regular sequences (with an average A, T, C, and G base content of 25% each). Important research areas, such as CGI regions (CpG islands), may require increased probe usage to obtain sufficient and accurate CGI data.
[0100] B. Amplicon-Based Sequencing Fragments of converted DNA may be amplified. In some embodiments, amplification is performed using primers designed to anneal to a methylation-converted target sequence having at least one methylation site. Methylation sequencing conversion converts unmethylated cytosines to uracil, while leaving 5-methylcytosines unaffected. A "converted target sequence" may also refer to a sequence in which cytosines known to be methylation sites are fixed as "C" (cytosine), while cytosines known to be unmethylated (unmethylated) are fixed as "U" (uracil; which may be treated as "T" (thymine) for primer design purposes).
[0101] In various examples, the source of DNA is cell-free DNA obtained from whole blood, plasma, or serum, or genomic DNA extracted from cells or tissues. In some embodiments, the amplified fragments are approximately 100-200 base pairs in length. In some embodiments, the DNA source is extracted from a cellular source (e.g., tissue, biopsy, cell line), and the amplified fragments are approximately 100-350 base pairs in length. In some embodiments, the amplified fragments contain at least one 20-base pair sequence containing at least one, at least two, at least three, or more than three CpG dinucleotides. Amplification may be performed using a set of primer oligonucleotides according to the present disclosure, and may use a thermostable polymerase. Amplification of multiple DNA segments may be performed simultaneously in one and the same reaction vessel. In some embodiments, two or more fragments are amplified simultaneously. For example, amplification may be performed using the polymerase chain reaction (PCR).
[0102] Primers designed to target such sequences may exhibit some bias toward converted methylated sequences. In some embodiments, PCR primers are designed to be methylation-specific for targeted methylation sequencing applications. Methylation-specific primers may enable higher sensitivity in some applications. For example, primers may be designed to contain a unique nucleotide (specific for methylated sequences after bisulfite conversion) positioned to achieve optimal discrimination, for example, in PCR applications. The unique nucleotide may be positioned at the 3'-end or penultimate position.
[0103] In some embodiments, primers are designed to amplify DNA fragments between 75 and 350 bp in length, a common size range for circulating DNA. Optimizing primer design based on target size can enhance the sensitivity of the methods described herein. Primers may also be designed to amplify regions of approximately 50 to 200, approximately 75 to 150, or approximately 100 or 125 bp.
[0104] In one embodiment, the amplifying step comprises using a primer that includes a unique dual index (UDI) sequence.
[0105] In one embodiment, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp in length.
[0106] In some embodiments, the methylation status of preselected CpG positions within a nucleic acid sequence is detected by an amplicon-based approach using methylation-specific PCR (MSP) primer oligonucleotides. The use of methylation-specific primers for the amplification of bisulfite-treated DNA allows for the discrimination of methylated and unmethylated nucleic acids. An MSP primer pair includes at least one primer that hybridizes to a converted CpG dinucleotide. Therefore, the primer sequence includes at least one CpG, TpG, or CpA dinucleotide. An MSP primer specific for unmethylated DNA includes a "T" at the 3' position of the C position in the CpG. Therefore, the base sequences of these primers may include sequences at least 18 nucleotides long that hybridize to the pre-treated nucleic acid sequence and its complementary sequence, and the base sequence includes at least one CpG, TpG, or CpA dinucleotide. In some embodiments of this method, the MSP primer includes two to five CpG, TpG, or CpA dinucleotides. In some embodiments, the dinucleotide is located within the 3' half of the primer; for example, in a primer having a length of 18 bases, the designated dinucleotide is located within the first 9 bases from the 3' end of the sequence. In addition to CpG, TpG, or CpA dinucleotides, the primer may further contain multiple methyl-converted bases (e.g., cytosine converted to thymine, or guanine converted to adenosine on the hybridized strand). In some embodiments, the primer is designed to have no more than two cytosine and / or guanine bases.
[0107] In some embodiments, each region is amplified in multiple portions using multiple primer pairs. In some embodiments, the portions do not overlap. The portions may be adjacent or spaced apart (e.g., spaced 10, 20, 30, 40, or 50 bp apart). Because target regions (including CpG islands, CpG shores, and / or CpG shelves) are typically longer than 75-150 bp, this example allows for assessment of the methylation status of sites across more (or all) of a given target region.
[0108] Primers can be designed for target regions using appropriate tools such as Primer3, Primer3Plus, or Primer-BLAST. As discussed, bisulfite conversion converts cytosine to uracil and 5'-methyl-cytosine to thymine. Therefore, primer positioning or targeting can utilize bisulfite-converted methylation sequences, depending on the degree of methylation specificity required.
[0109] III. Library preparation for enzymatic methylation sequencing In the first aspect, a method for preparing a sequencing library is provided.The method described herein provides a library that is acceptable for both next-generation unmethylated and methylated sequencing applications, thereby providing the sequencing data for two applications from a single sample.The obtained raw sequencing data can be used not only for analyzing methylation status, but also for traditional cfDNA analysis such as copy number changes, germline variant detection, somatic variant detection, nucleosome positioning, transcription factor profiling, chromatin immunoprecipitation, etc.
[0110] A. Adapter Ligation for Targeted Sequencing Applications In one embodiment, the method of the present invention preserves the integrity and information of nucleic acid sequences for methylation profiling. In one example, combining dsDNA adapter ligation prior to enzymatic conversion preserves the endpoint information of fragments while providing the highest possible library complexity for target enrichment (or directly for genome-wide sequencing), thereby improving the sensitivity of detecting rare events such as methylated ctDNA. The advantages and comparison of adapter ligation prior to conversion are shown in Figure 1.
[0111] In one example, nucleic acid adapters are ligated to the 5' and 3' ends of a population of nucleic acid fragments in a biological sample to generate a sequencing library. In one example, a collection of nucleic acid adapters is ligated to the nucleic acid fragments in the sample, where the collection of adapters contains equal portions of 4-bp, 5-bp, and 6-bp unique molecular identifier (UMI) sequences to enable T / A overhang ligation, followed by an invariant thymidine (T) at the final position (i.e., the 3' end). In this way, the UMI is located adjacent to the insert nucleic acid of the library. During sequencing, the UMI is also sequenced as part of the read at the 5' end (alternatively, the UMI matches the library insert at the sequencing read level). The invariant T is staggered across three positions to maintain base diversity at the sequencing position. In contrast, using a single-length UMI with an invariant thymidine results in low-complexity sequencing at the position corresponding to the invariant thymidine, resulting in reduced sequencing quality. The first 4 bp of each UMI contains a set of 4-bp core UMI sequences with an edit distance of 2 or greater and balanced nucleotides and colors. Despite the variable length of UMI sequences, the use of a single-length core UMI facilitates the use of bioinformatics tools built for single-length UMIs for UMI extraction and deduplication. In this way, the 4-bp core sequence serves as a recognition sequence that informs bioinformatics tools to trim 5, 6, or 7 bases (including the invariant T), thereby maintaining accurate cfDNA endpoint information. A schematic illustrating staggered adapters is shown in Figure 2. The use of UMIs enables read deduplication, single-strand error correction, and duplex reconstruction after sequencing, thereby enabling the use of reverse complements of reads to enhance error correction, also known as double-strand error correction. In another example, unique dual indexes (UDIs) are additional sequences that can be added to UMI-containing adapters during library purification to provide sample barcoding and demultiplexing of samples after sequencing.In various examples, the UDI sequence is 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, or 12 bp in length.
[0112] In various embodiments, the nucleic acid adapter may comprise a 4-6 bp long UMI with a 5' thymidine overhang, which is designed to be non-unique (i.e., drawn from a specific, constrained set of sequences).
[0113] In one embodiment, some UMIs contain one or more methylcytosine bases. The efficiency of enzymatic methylation conversion reactions (including TET oxidation and APOBEC deamination) can be evaluated by the UMI mismatch rate, based on the proportion of UMIs that do not match a specific constrained set of designed UMI sequences. The UMI mismatch rate can be used as an embedded quality control indicator to evaluate the quality of a sequencing library. In addition, when a perfect UMI match is required in a bioinformatics pipeline, the UMI mismatch rate can be used as a filter to remove individual reads that may be of low quality due to incomplete conversion.
[0114] In various embodiments, the UMI mismatch rate is less than 6%, less than 5%, less than 4%, less than 3%, or less than 2%.
[0115] In another embodiment, the UMI contains one or more cytosines with modifications that can be used to monitor enzymatic activity. Non-limiting examples of these modified bases include 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine.
[0116] B. Enzymatic Conversion for DNA Methylation Sequencing Applications Tet-assisted pyridine borane sequencing (TAPS) is a minimally destructive methylation sequencing method for converting cytosine to uracil in nucleic acids. This bisulfite-free method minimizes DNA degradation, achieving conversion rates comparable to sodium bisulfite sequencing while preserving the length of nucleic acid molecules. TAPS results in higher sequencing quality scores for cytosine and guanine base pairs and more uniform coverage of various genomic features, such as CpG islands.
[0117] In TAPS, the ten-eleven translocation (Tet1) enzyme oxidizes both 5mC and 5hmC to 5caC. Pyridine borane reduces 5caC to dihydrouracil, a uracil derivative that is converted to thymine after PCR. TAPS can be performed in two other ways: TAPSβ and chemically assisted pyridine borane sequencing (CAPS). In TAPSβ, β-glucosyltransferase is used to label 5hmC with glucose, protecting it from oxidation-reduction reactions and allowing for specific detection of 5mC. In CAPS, potassium perruthenate acts as a chemical surrogate for Tet1 and specifically oxidizes 5hmC, allowing for direct detection.
[0118] In one example, the combination of enzymatic conversion of unmodified C to U and staggering UMI adapters to match library inserts is useful for targeted sequencing of methylated libraries. This combination may allow for a reduced plasma quantitative input or a reduced cfDNA mass input compared to bisulfite conversion sequencing, since sample cfDNA is not degraded to the same extent in low-depth sequencing applications.
[0119] For high-depth sequencing applications, cfDNA is not degraded to the same extent, potentially resulting in higher depth sequencing compared to bisulfite conversion sequencing from similar inputs of plasma or cfDNA.
[0120] In one example, cytosines present in the adapter nucleic acid are modified with 5-methyl or 5-hydroxymethyl groups to prevent CT conversion in the adapter.
[0121] One advantage of this approach is that adapter ligation prior to conversion preserves fragment endpoint and length information compared to bisulfite conversion followed by ssDNA adapter ligation, where significant degradation of the nucleic acid prior to adapter ligation can result in the loss of informative fragment endpoint and length information.
[0122] Enzymatic conversion of unmodified C to U is less stressful on sample nucleic acid fragments and may result in more complete and uniform coverage compared to bisulfite methods. Bisulfite degradation of DNA is not uniform, with some sequences preferentially degraded over others, such as CG dinucleotides, the very sites interrogated by methylation sequencing. Therefore, enzymatic approaches yield higher coverage of CpG sites than bisulfite conversion methods with the same number of unique reads, resulting in higher uniformity of captured reads in target enrichment applications. Furthermore, non-bisulfite methods (e.g., enzymatic and chemical conversion methods like TAPS) offer increased resolution of biological signals, specifically the ability to distinguish between 5mC and 5hmC methylation in nucleic acid sequences. This information and additional resolution can be beneficial in computational approaches and other methods.
[0123] In some examples, exposing DNA or barcoded DNA to an enzymatic reaction that converts cytosine nucleobases of the DNA or barcoded DNA to uracil nucleobases comprises "performing an enzymatic conversion."
[0124] In various examples, glycosylation and oxidation reactions overcome the observed inherent deamination of 5hmC and 5mC by deaminases. Deaminases convert 5mC and unmodified C to U, but not 5ghmC and 5caC. Non-limiting examples of deaminases include APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like). The embodiments described herein utilize enzymes with substantially no sequence bias in glycosylation, oxidation, and deamination of cytosine. Furthermore, these embodiments do not substantially impart nonspecific damage to DNA during glycosylation, oxidation, and deamination reactions.
[0125] In some embodiments, a glucosyltransferase (GT), e.g., β-glucosyltransferase (βGT), is used to covalently attach glucose to 5hmC and protect this modified base from deamination. Other enzymatic or chemical reactions to modify 5hmC can also be used to achieve the same effect.
[0126] In general, and in one aspect, the methods provided herein include (a) treating an aliquot (portion) of a nucleic acid sample with a dioxygenase, e.g., TET2, and βGT in a reaction mixture to generate a reaction product in which substantially all modified cytosines (C) are oxidized or, in the case of 5hmC, glucosylated; and (b) treating this reaction product with cytidine deaminase to convert substantially all unmodified Cs to U. The term "modified" cytosine used throughout these examples and embodiments refers to one or more of 5mC, 5hmC, 5ghmC, 5fC, and 5caC, where oxidation of 5mC, 5hmC, and 5fC to completion produces 5caC. βGT reacts only with 5hmC. However, some 5hmC may be converted to 5fC by a dioxygenase and further converted to 5caC before glucosylation occurs. In the presence of dioxygenases, 5mC is mostly oxidized to 5caC, but some residual 5hmC may be generated. However, the residual 5hmC can be glycosylated by βGT to prevent the low deamination rate of 5hmC, which would otherwise reduce the accuracy of methylation sequencing.
[0127] Therefore, the described method largely distinguishes between unmodified and modified cytosines by treating nucleic acids with dioxygenase before deamination. However, the amount of naturally occurring 5mC in genomic DNA can substantially exceed the amount of 5hmC, and therefore can exceed the amount of naturally occurring 5fC and 5caC. Therefore, the amount of naturally occurring modified cytosine is generally considered to be an approximation of the amount of naturally occurring 5mC.
[0128] In one embodiment, the method is adaptable for performing 5hmC sequencing. The 5hmC sequencing method further comprises treating an aliquot of a nucleic acid sample with βGT in the absence of dioxygenase, followed by treatment with cytidine deaminase, to produce a reaction product in which substantially all 5hmC in the aliquot is glucosylated and substantially all unmodified C and 5mC are converted to U. After PCR amplification, U is converted to T, and therefore cytosine and 5mC are indistinguishable upon sequencing. The resulting reaction product is sequenced and compared to a reference sequence, allowing 5hmC to be distinguished from C and 5mC. By distinguishing these sites, these modified nucleotides can be mapped to a reference sequence, e.g., a reference sequence from a database or an independently determined reference sequence.
[0129] In some embodiments, the reaction product of a dioxygenase and a deaminase with βGT, or its amplification product, can be sequenced to determine which Cs are methylated (which may contain a small fraction of 5hmC) and which Cs are unmodified. In some embodiments, the reaction product of a deaminase and a βGT without a dioxygenase, or its amplification product, can be sequenced to determine which Cs are hydroxymethylated and which Cs are unhydroxymethylated. In some embodiments, the reaction product of a deaminase and a βGT without a dioxygenase, or its amplification product, can be sequenced to determine which Cs are hydroxymethylated and which Cs are unmodified. Reference DNA can be generated by sequencing the reaction products produced by reacting a nucleic acid sample with no one of dioxygenase, βGT, and deaminase. Alternatively, the reference sequence can be a known reference sequence, for example, from a sequence database.
[0130] In one embodiment, the sequence of the reaction product of a dioxygenase and a deaminase with βGT can be compared to a reference sequence, which can optionally be compared to the sequence of the reaction product of βGT (without dioxygenase) and a deaminase to determine which cytosines in a nucleic acid sample are modified with methyl groups versus hydroxymethyl groups.
[0131] In one aspect, a method for targeted methylation sequencing of a nucleic acid sample is provided, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to cfDNA, wherein the cfDNA comprises unconverted nucleic acid; b) enzymatically converting unmethylated cytosines to uracils within the nucleic acid molecule to produce converted nucleic acids; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with nucleic acid probes complementary to a panel of pre-identified CpG or CH loci to enrich for sequences corresponding to said panel of pre-identified CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid to reference nucleic acid sequences of a panel of pre-identified CpG or CH loci to determine a methylation profile of a cell-free DNA (cfDNA) sample from the subject.
[0132] If the test converted nucleic acid sequence is a T corresponding to the reference C at the designated CpG locus, then the C was unmethylated in the original test nucleic acid fragment. In contrast, if the test converted nucleic acid sequence and the reference sequence are both a C at the designated CpG locus, then the C was methylated in the original test nucleic acid fragment.
[0133] In one example, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 50-500x, about 25-1000x, about 50-500x, about 250-750x, about 500-200x, about 750-1500x, or about 100-2000x. In some embodiments, the nucleic acid sequence is sequenced at a depth of 100x or greater than 500x.
[0134] In one example, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 500x, about 1000x, about 2000x, about 3000x, about 4000x, about 5000x, about 6000x, about 7000x, about 8000x, about 9000x, about 10000x, or greater than 5000x.
[0135] In one example, the nucleic acid sequence of the converted nucleic acid molecule is sequenced to a depth of about 300-fold unique, about 400-fold unique, about 500-fold unique, about 600-fold unique, about 700-fold unique, about 800-fold unique, about 900-fold unique, or about 1000-fold unique, or greater than 500-fold unique.
[0136] C. Target Enrichment Sequencing Applications Additionally, a method for enriching desired methylated regions in target capture applications during sequencing is provided. A potential problem with applying target enrichment capture panels using DNA methylation libraries is the low rate of on-target reads and the high rate of off-target DNA fragment capture. For each region in the panel, probes can be designed to target DNA derived from methylated CpGs or DNA derived from unmethylated CpGs. For either probe type, all CpG sites along the region are considered unmethylated or methylated, as appropriate for the probe type. The probes can be hybridized to library molecules after bisulfite / enzyme conversion and PCR amplification. Only the library molecules captured by the probes are then sequenced. This method has the advantage of reducing sequencing costs because only a small portion of the genome is sequenced. In one example, approximately 0.1% of the genome is sequenced. In one example, approximately 0.3% of the genome is sequenced. In one example, approximately 0.5% of the genome is sequenced. In one example, approximately 0.7% of the genome is sequenced. In one example, approximately 1% of the genome is sequenced. In other examples, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the genome is sequenced.
[0137] Both bisulfite and enzyme-converted library target capture enrichment approaches can result in significant off-target capture rates. Off-target capture rates are due in part to the conversion of all cytosines not at CpG sites to T in both types of probes hybridizing to DNA derived from methylated CpGs. The reduction in cytosine content in the probes leads to a decrease in sequence complexity and therefore reduces the specificity of the probes hybridizing to the target library molecules.
[0138] As used herein, the terms "methylated probe" and "unmethylated probe" refer to probes used to hybridize to methylated CpGs and unmethylated CpGs, respectively, in a converted nucleic acid sequence. The probes may be designed to recognize the converted nucleic acid sequence. In a converted methylated CpG probe, C remains as C after conversion. In a converted unmethylated CpG probe, C is converted to T after conversion. In both the converted methylated probe and the unmethylated probe, all Cs in non-CpG dinucleotides are converted to T after conversion.
[0139] Methylated probes retain some cytosines (i.e., cytosines at CpG sites). In contrast, in unmethylated probes, all cytosines are converted to thymine. Unmethylated probes are less complex than methylated probes and are more likely to contribute preferentially to off-target capture rates. In one example, a probe that hybridizes to DNA derived from methylated CpG is used in a target enrichment method. In one example, a probe that has a sequence substantially complementary to a target that hybridizes to DNA derived from methylated CpG is used in a target enrichment method.
[0140] For target enrichment, the probes that hybridize to DNA derived from methylated CpG can be selected to realize different embodiments. The target capture hybridization reaction occurs at a single temperature. However, the optimal melting temperature (Tm) of the probes that hybridize to DNA derived from methylated CpG is, on average, higher than the Tm of the probes that are not designed to hybridize to DNA derived from methylated CpG.
[0141] A cytosine base pair contains three hydrogen bonds, while a thymine base pair contains two. Converting a cytosine to a thymine in a probe reduces the hydrogen bonds, thereby decreasing the probe's Tm. Because methylated probes contain some cytosines, whereas unmethylated probes retain none, the methylated probe has a higher Tm than its matched unmethylated counterpart. As the number of CpG sites in a region increases, the difference in melting temperature between the methylated and unmethylated probes also increases. Probes with higher melting temperatures may hybridize to target DNA fragments more efficiently than probes with lower melting temperatures. Generally, hybridization temperatures are selected to be relatively high to facilitate on-target capture. However, at typical hybridization temperatures, methylated probes hybridize more efficiently than unmethylated probes due to the higher melting temperature resulting from the retention of multiple cytosines. High melting temperatures can bias the % CpG methylation levels measured by the target capture hybridization approach compared to levels measured by sequencing the pre-capture library.
[0142] In one example, only a single probe type, either methylated or unmethylated, is used in the hybridization reaction to enrich for hypermethylated or hypomethylated library molecules, respectively. Using a single probe type, either methylated or unmethylated, avoids the issue of different melting temperatures between probe types. Using a single probe type also allows for more efficient capture (or enrichment) of the same DNA fragment type. In one example, using only methylated probes results in preferential binding of hypermethylated ROIs over hypomethylated ROIs. In another example, using only unmethylated probes results in enrichment of unmethylated ROIs.
[0143] Using only a single probe type also allows for the use of higher hybridization temperatures to reduce off-target capture without affecting the relative balance of capture of methylated and unmethylated ROIs. Thus, probe panels can be designed based on the desire to enrich for hypermethylated or hypomethylated DNA fragments. As an example, if quantification of both hypermethylated and hypomethylated DNA fragments is desired, two parallel but separate hybridization reactions are employed for both methylation states.
[0144] D. Methylation Analysis In various examples, once enzymatic methylation sequencing is completed, the assay is used to analyze the methylation status of nucleic acids in a biological sample. In one example, whole-genome enzymatic methyl-sequencing ("WG EM-seq") provides high-resolution sequencing by characterizing DNA methylation at nearly every cytidine nucleotide in the genome. Other targeted methods, such as targeted enzymatic methyl-sequencing ("TEM-seq"), can be useful for methylation analysis.
[0145] In other examples, assays traditionally used for bisulfite conversion can be adapted to minimally disruptive conversion methods such as enzymatic conversion, TAPS, and CAPS. In various examples, assays used for methylation analysis can be mass spectrometry, methylation-specific PCR (MSP), reduced expression bisulfite sequencing (RRBS), HELP assay, GLAD-PCR assay, ChIP-on-chip assay, restriction landmark genome scanning, methylated DNA immunoprecipitation (MeDIP), pyrosequencing of bisulfite-treated DNA, molecular break light assay, methyl-sensitive Southern blotting, high-resolution lysis (HRM or HRMA), ancient DNA methylation reconstruction, or methylation-sensitive single-nucleotide primer extension assay (msSNuPE).
[0146] The methylation profile of cfDNA can be identified by applying sequencing alignment methods to map methyl sequence reads obtained from whole-genome or targeted methyl sequencing of a human reference genome. Non-limiting examples of sequence alignment methods include bwa-meth, bismark, Last, GSNAP, BSMAP, NovoAlign, Bison, metagenomic phylogenetic analysis (e.g., MetaPhlAn2), BLAT, Burrows-Wheeler Aligner (BWA), Bowtie, Bowtie2, Bfast, BioScope, CLC bio, Cloudburst, Eland / Eland2, GenomeMapper, GnuMap, Karma, MAQ, MOM, Mosaik, MrFAST / MrsFAST, PASS, PerM, RazerS, RMAP, SSAHA2, Segemehl, SeqMap, SHRiMP, Slider / SliderII, Srprism, Stampy, vmatch, ZOOM, and SOAP / SOAP2 alignment tools.
[0147] E. CpG Error Correction Using Duplex UMI-Based Methylation Consensus Calling Methylation analysis involves analyzing sequencing data based on whether a "C" in a CpG context is read as a "C" (methylated) or a "T" (unmethylated) in the sequencing. However, "T" can appear at these positions in the sequencing for reasons other than the presence of an unmethylated CpG in the parent DNA molecule. These reasons include sequencing errors, PCR errors, fill-in nucleotides during end repair, DNA damage, germline single nucleotide polymorphisms (SNPs) that replace a CpG with another dinucleotide, somatic mutations that replace a CpG with another dinucleotide, and overconversion (i.e., the conversion of a 5mC to a T despite the presence of a methylation mark). In addition, "C" can appear at these positions in the sequencing for reasons other than the presence of a methylated CpG in the parent DNA molecule. These reasons include sequencing errors, PCR errors, DNA damage, and incomplete conversion (the failure of an unmethylated C to convert to a T despite the absence of a methylation mark). Failure to correct for most of these error modes would result in inaccurate readings of CpG methylation status, potentially limiting the detection of rare events (e.g., ctDNA molecules in early-stage cancers) that require highly accurate readings of CpG methylation status. Additionally, methods that cannot account for duplex structure information cannot distinguish between hemimethylated and symmetrically methylated CpG sites. Such information is useful for interpreting the biological significance of methylation signals. For example, hemimethylation can directly identify de novo methylation events, thereby distinguishing between de novo and maintenance factors.
[0148] Duplex sequencing approaches overcome the limitations of sequencing accuracy by addressing these widespread errors. For example, duplex sequencing reduces errors by independently tagging and sequencing each of the two strands of a DNA duplex. Because the two strands are complementary, true mutations can be found at the same position on both strands. Similarly, because CpG dinucleotides are symmetric, fully methylated CpG motifs have methylated cytosines at opposing adjacent positions on both strands. In contrast, PCR or sequencing errors occur on only one strand. This method uniquely exploits the redundant and additional information present between the strands of double-stranded DNA, thereby overcoming the technical limitations of methods that utilize data from a single strand.
[0149] For enzymatic methylation sequencing, the efficiency of APOBEC conversion of individual fragments can be assessed by the number of cytosines in the CHH context that are sequenced as cytosine. In a 100% efficient APOBEC reaction, all cytosines in the CHH context are converted to uracil and sequenced as thymine. In contrast, cfDNA fragments in which the APOBEC enzyme did not work efficiently (i.e., conversion was incomplete) may contain one or more cytosines in the CHH context that were not converted to uracil, which will be sequenced as cytosine. The number of unconverted cytosines in the CHH context can be used as a filter to remove reads that may be unreliable and noisy due to incomplete conversion.
[0150] Many enzymes that act on nucleic acids have sequence preferences that result in biases in which sites are efficiently acted upon by the enzyme. Experimental data can be used to identify the sequence preferences of individual enzymes. In various embodiments, this data can be used to mask potential sites that are likely to be incompletely converted by the enzyme. As an example, APOBEC A3A has a 12-fold higher discrimination against cytosines preceded by an A compared to cytosines preceded by a T.
[0151] In one example, a method for double strand methylation consensus calling is provided, the method comprising: a) preparing a methylation sequencing library from cfDNA using enzymatic conversion, (i) ligating double-stranded adapters to nucleic acid fragments obtained from a biological sample; (ii) performing target enrichment to enable ultra-deep sequencing of pre-identified desired loci; (iii) preparing methylation sequencing libraries from cfDNA using enzymatic conversion (so that neither strand is damaged) and ligating duplex UMIs prior to enzymatic conversion (to tag the duplex prior to the denaturation step involved in enzymatic conversion); and b) a target enrichment step that allows for ultra-deep sequencing of specific loci of interest; c) sequencing the enriched library using single-end or paired-end reads; d) correcting sequencing errors within the overlapping region of the paired-end reads for the sequencing fragments of the paired-end reads; e) collapsing the sequencing fragments into concatenated read families to correct errors resulting from PCR and sequencing; f) collapsing concatenated read families into double read families to identify discrepancies in the putative methylation status of symmetric CpGs.
[0152] A schematic diagram of this method is shown in FIG.
[0153] CpG "error correction" of methyl-sequence data using duplex information offers the advantage of filtering noise that could otherwise reduce the sensitivity or specificity of classifiers using methyl-sequence data as input. Because nucleotide imbalances are introduced into sequences after conversion, unique UMI designs using staggered methylated UMIs in the context of methyl-sequencing may increase data output by improving sequencing accuracy and aiding in cluster identification (especially using platforms such as the NextSeq sequencer), reducing reliance on adding large amounts of PhiX data (increasing sequencing depth and reducing associated costs). Unlike standard duplex sequencing, which analyzes the concordance of variant calls at base-paired nucleotides, CpG duplex-based duplex sequencing methods evaluate the symmetry (1-bp offset) of CpG methylation between strands. In certain instances, duplex sequencing can also distinguish SNPs from unmethylated CpGs. Enzymatic methyl-sequencing methods have the advantage over bisulfite-based methods of efficiently capturing sequences on both strands.
[0154] F. CpG Error Correction Using Conversion-Resistant Adapters for Duplex UMI-Based Methylation Consensus Calling in Enzymatic Methylation Sequencing In another embodiment, conversion-resistant adapters and primers are used for methyl sequencing. Sequencing methods used to identify the location of base modifications, such as bisulfite sequencing or enzymatic methylation sequencing (EM-seq), used to identify 5mC, operate by chemically or enzymatically modifying each unmodified cytosine base (C) to change the base-pairing properties of C. For example, during the EM-seq process, all unmodified Cs are converted to uracil (U) by APOBEC enzymes and then sequenced as thymine (T). 5mC bases are not converted and are sequenced as C. Because bases can only be converted when DNA is single-stranded, double-stranded DNA must be denatured before the C-to-U conversion reaction.
[0155] One potential problem when combining duplex sequencing and methylation sequencing is reduced PCR amplification and sequencing. Because adapters must be ligated onto DNA while it is still double-stranded (i.e., before base conversion), all Cs in the adapters will be converted to Us, thereby preventing efficient PCR amplification and sequencing.
[0156] A solution to this problem is to use adapters with modified bases (e.g., 5mC, 5hmC, or other C variants) that are not or are poorly converted during the deamination reaction. However, oligonucleotides containing modified bases are often significantly more expensive than oligonucleotides containing only standard bases. Furthermore, this solution is generally only effective for bisulfite methyl sequencing, where 5mC is not unconvertible.
[0157] Unlike bisulfite sequencing, the EM-seq process requires an additional enzymatic step to prevent APOBEC-mediated conversion of 5mC or 5hmC to U. This step uses either the Tet enzyme, which oxidizes the 5mC or 5hmC base, or βGT, which glycosylates 5hmC, thereby protecting 5hmC from conversion. If this step is not completely efficient, some of the 5mC or 5hmC in the adapters will be converted to uracil, resulting in loss of library complexity and reduced sequencing quality. The Tet oxidation reaction is highly sensitive to reaction conditions, which can lead to variable sequencing library quality.
[0158] To improve the oxidation robustness of EM-seq duplex sequencing (and reduce its current economic burden), conversion-resistant adapters containing only unmodified bases can be used. Unmodified bases refer to the traditional bases guanine, cytosine, adenine, and thymine in their unmodified state. Contrary to traditional methods that limit total conversion of adapter molecules using modified bases such as 5mC and 5hmC, this approach allows all cytosine conversions in the adapter to improve efficiency and sequencing quality. An example of a conversion-resistant adapter is shown in Figure 4, Panel A.
[0159] Sequencing libraries generated using these converted resistant adapters can be amplified and sequenced using a set of PCR and sequencing primers that match the original adapter sequences. After conversion, sequencing libraries can be amplified and sequenced using PCR and sequencing primers that match the converted adapter sequences in Figure 4, panel B.
[0160] G. Use of Internal Process Control During Enzymatic Conversion For targeted enzymatic methylation sequencing, synthetic internal process control (IPC) may be used to monitor the oxidation and deamination reactions during enzymatic methylation conversion.
[0161] In various embodiments, the IPC may include all 256 possible cytosine contexts in a window of two bases before and after the C (NNCNN).
[0162] In various embodiments, the IPC is a duplex synthesized by PCR containing either 100% unmodified C, 100% methylated C, or 100% hydroxylated C (or another modification to C). In this regard, the conversion or protection efficiency of the IPC can be monitored. In some embodiments, the conversion or protection efficiency can be monitored by sequencing or quantitative PCR.
[0163] H. Hemimethylation Analysis In another example, the use of UMIs in methyl sequences allows for error correction and analysis / removal of hemimethylation. Alternatively, strand-specific methylation sequencing allows for the identification of hemimethylated DNA. The methylation status of CpG / CpG dyads is typically concordant, i.e., fully methylated or completely unmethylated. However, CpG / CpG dyads with discordant methylation status, i.e., hemimethylated, generally occur at low to moderate frequency, except in regions of transcriptional silencing or reactivation, or when they arise transiently during DNA replication. Such hemimethylated dyads provide additional information that can inform classifiers when stratifying populations. Recognizing hemimethylated dyads allows for a more complete methyl sequence profile, allowing for the selection of whether to remove or include this information when generating classifiers.
[0164] Another advantage of enzymatic methyl-sequencing approaches is their ability to better distinguish between methylated and unconverted Cs. By maintaining fragment integrity and length through enzymatic conversion, duplex UMI methyl sequences can be used to increase the accuracy of determining the true methylation status of nucleic acid molecules. This method can account for errors that may arise during, for example, extraction (DNA damage), library preparation (end repair fill-in), enzymatic conversion (under- or over-conversion), PCR (base incorporation errors), and sequencing (base calling errors). Increasing the accuracy of methylation status determination improves characterization and the generation of classifiers for stratifying populations using these methylation-based epigenetic sequence differences. In one example, adapter orientation is used to distinguish dsDNA fragments originating from the top strand versus the bottom strand (based on which genomic strand read1 maps), as shown schematically in Figure 3. This method contrasts with methods that rely on index barcodes for error correction.
[0165] I. Identification of somatic mutants In various examples, enzyme-converted DNA is used to infer the methylation status of C residues in the genome. However, because enzymatic conversion of DNA converts unmethylated C residues to U residues without introducing other chemical changes into the DNA, somatic variants that do not correspond to C or T bases in the reference or query sequence can also be identified in the converted DNA. These somatic variants can be identified using existing methods designed for unconverted DNA (including error-correction methods such as duplex sequencing). Furthermore, somatic variants that correspond to C or T bases in the reference or query sequence can be distinguished from methylation-associated sequencing patterns using duplex sequencing, based on the expectation that somatic variants should be found at the same position on both strands of a duplex DNA molecule, whereas methylation-associated patterns do not (i.e., because C and T bases do not base-pair with each other). This difference allows EM sequencing to identify both the methylation status and somatic variants at CpG sites.
[0166] J. Estimation of Nucleosome Position Cytosine methylation at CpG sites can be significantly enriched in DNA spanning nucleosomes compared to adjacent DNA. Therefore, CpG methylation patterns can be further employed to infer nucleosome positions using machine learning approaches. EM sequence datasets, regardless of methylation status, can be analyzed according to the same methods used for WGS to generate features that are input into machine learning methods and models. The 5mC patterns can then be used to predict nucleosome positions, which can be useful for inferring gene expression and / or classifying diseases and cancers. In another example, features can be derived from a combination of methylation status and nucleosome position information.
[0167] Metrics used in methylation analysis include, but are not limited to, M-bias (base-wise methylation %) of CpGs, CHGs, and CHHs, conversion efficiency (e.g., %100 average methylation for CHHs), hypomethylated blocks, methylation level (e.g., overall average methylation of CPGs, CHHs, CHGs, chrMs, LINE1, or ALUs), dinucleotide coverage (normalized coverage of dinucleotides), coverage evenness (e.g., unique CpG sites at 1x and 10x average genome coverage (when running S4), overall average CpG coverage (depth), and average coverage at CpG islands, CGI shelves, and CGI shores. In one example, the output of duplex-based CpG methylation calls is used as input for this analysis. In one example, fragment endpoint and length information is used as feature input for the analysis. These metrics may be used as feature input for machine learning methods and models.
[0168] In another aspect, the disclosure provides a method, the method comprising: (a) providing a biological sample containing cfDNA from a subject; (b) exposing the cfDNA to conditions sufficient to optionally enrich for methylated cfDNA in the sample; (c) enzymatically converting unmethylated cytosine nucleobases of the cfDNA to uracil nucleobases; (d) sequencing the cfDNA, thereby generating sequence reads; and (e) computationally processing the sequence reads to (i) determine the degree of methylation of the cfDNA based on the presence of uracil nucleobases, and (ii) model at least partial degradation of the cfDNA, thereby generating degradation parameters; and (f) using the degradation parameters and the degree of methylation to determine gene sequence characteristics.
[0169] In some examples, sequencing cfDNA includes determining the degree of DNA methylation based on the ratio of unconverted cytosine nucleobases to converted cytosine nucleobases. In some examples, converted cytosine nucleobases are detected as uracil nucleobases. In some examples, uracil nucleobases are observed as thymine nucleobases in sequence reads.
[0170] In some examples, generating the resolution parameters includes using a Bayesian model. In some examples, the Bayesian model is based on strand bias or enzymatic conversion or over-conversion. In some examples, computing the sequence reads includes using the resolution parameters under the framework of a pairwise HMM or a naive Bayesian model.
[0171] K. Analysis of variably methylated regions (DMRs) In one example, the methylation analysis is a variably methylated region (DMR) analysis. DMRs are used to quantify CpG methylation across regions of the genome. Regions are dynamically assigned through discovery. Many samples of different classes can be analyzed to identify regions with the most variably methylated regions across various classifications. A subset of regions can be selected to have variably methylated regions and used for classification. The number of CpGs captured in a region may also be used in the analysis. In one example, the output of duplex-based CpG methylation calling is used as input for this analysis. Regions may be variable in size. In one example, a pre-discovery process is performed that bundles many CpG sites together as regions. In one example, DMRs are used as input features for machine learning methods and models.
[0172] L. Methylation haplotype blocks and methylation haplotype loads In one example, a haplotype block assay is applied to samples. Identification of methylation haplotype blocks aids in the deconvolution of heterogeneous tissue samples and tumor tissue origin mapping from plasma DNA. Tightly connected CpG sites, known as methylation haplotype blocks (MHBs), can be identified in WGBS data. To perform tissue-specific methylation analysis at the block level, a metric called methylation haplotype load (MHL) is used. This method provides informative blocks useful for deconvolution of heterogeneous samples. This method is useful for quantitative estimation of tumor burden and tissue origin mapping in circulating cfDNA. In one example, the output of duplex-based CpG methylation calling is used as input for this analysis. In one example, haplotype blocks are used as input features for machine learning methods and models.
[0173] M. Targeted Methylation Calling Analysis for Identifying Cell-Type of Origin In one embodiment, a method for target methylation calling is used to identify the cell type origin of cfDNA molecules based on methylation pattern.This method provides a probabilistic model of the joint methylation status of multiple adjacent CpG sites on each sequencing read, so as to take advantage of the global nature of DNA methylation for signal amplification.This model develops the probability of the sequencing read of each cell type, and then develops a comprehensive cell type mixture model and fits this model.
[0174] Traditional DNA methylation analysis focuses on the methylation rate (β value) of individual CpG sites in a cell population, indicating the proportion of cells in which that CpG site is methylated. Such population-averaged measurements are often insufficiently sensitive to capture aberrant methylation signals that affect only a portion of the cfDNA. However, based on the widespread nature of DNA methylation, disease-specific cfDNA reads can be computationally distinguished from normal cfDNA reads.
[0175] Additionally, given the widespread nature of DNA methylation, the co-methylation status of multiple adjacent CpG sites can be used to easily distinguish cancer-specific cfDNA reads from normal cfDNA reads. The average methylation value (referred to as the α value) of all CpG sites in a given read provides the difference (0 and 1) between aberrantly methylated cfDNA and normal cfDNA (α tumor = 0%, α normal = 100%). The methylation α value is used to estimate whether the combined probability of all CpG sites in a read follows the DNA methylation signature of the disease. This method can sensitively identify cfDNA from multiple cell types among all cfDNA in plasma.
[0176] In various examples, alignment tools are used to align reads to a reference genome and call methylated cytosines. PCR duplicates are removed, and the number of methylated and unmethylated cytosines is quantified for each CpG site. The methylation level of a CpG cluster is calculated as the ratio of the number of methylated cytosines to the total number of cytosines in the cluster. This WGBS data processing procedure calculates the average methylation level of a CpG cluster in normal plasma samples, which is used to identify methylation markers. When plasma cfDNA samples are used as test data, the combined methylation status of all CpG sites in individual sequencing reads aligned to the marker panel regions is extracted and input into a machine learning model. In this method, duplex-based CpG methylation calls are used as input features for methylation status analysis and feature generation. To improve the input data quality of high-coverage cfDNA methylation data, reads covering more than two, three, or four CpG sites can be filtered.
[0177] The methylation sequencing method described herein can improve the quality of sequencing reads, for example, by reducing PCR error and bias, and reducing the DNA degradation that occurs in bisulfite conversion.In one example, methylation sequencing data is used to model overlapping regions.In one example, machine learning modeling can determine the cell type origin for identified methylated DNA regions.
[0178] In various examples, the model can classify sequences from two or more cell types, hi other examples, the model can classify sequences into 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 50, 75, 100, or more than 100 different cell types.
[0179] N. DNA hydroxymethylation analysis In one embodiment of the present invention, 5hmC sequencing can be achieved by substituting hydroxymethylation in the adapter nucleic acid in the adapter ligation step, and then using only βGT to attach glucose to 5hmC residues in the test nucleic acid library insert, instead of using dioxygenase and βGT to attach 5mC and 5hmC. When the resulting sequencing data is compared to a reference genome, all C positions in the reference that show a corresponding C in the test sequence are interpreted as hydroxymethylated C, and all Cs in the reference that show a T in the test sequence are interpreted as unmodified C or methylated C. Therefore, data interpretation for hydroxymethylation analysis is the same as for methylation analysis.
[0180] In one embodiment of the invention, methylated and hydroxymethylated sequencing libraries can be compared to identify the level of each cytosine modification (e.g., 5m or 5mC) at single nucleotide resolution.
[0181] In one aspect of the invention, all analytical methods used with methylation sequencing data can be applied to hydroxymethylation sequencing data because the readout of hydroxymethylation status is the same as the methylation status.
[0182] IV. Computer Systems and Machine Learning Methods A. Sample characteristics As used herein, as it relates to machine learning and pattern recognition, the term "feature" can also refer to individual measurable characteristics or traits of an observed phenomenon. Features are typically numeric, although structural features such as strings and graphs are used in syntactic pattern recognition. The concept of a "feature" is related to the concept of an explanatory variable used in statistical methods such as linear regression.
[0183] In one embodiment, the features are input into a feature matrix for machine learning analysis.
[0184] For multiple assays, the system identifies a feature set as input to a machine learning model. The system runs an assay for each molecular class and forms a feature vector from the measurements. The system inputs the feature vector into the machine learning model to obtain an output classification of whether the biological sample has the specified characteristic.
[0185] In one embodiment, the machine learning model outputs a classifier that distinguishes between two groups or classes of individuals, i.e., a characteristic in a population or a characteristic of a population of individuals. In one embodiment, the classifier is a trained machine learning classifier.
[0186] In one embodiment, the information-rich loci or features of biomarkers are assayed in cancer tissue to form a profile.Receiver operating characteristic (ROC) curve is useful for plotting the performance of a specific feature (for example, any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing two groups (for example, the individuals who respond to therapeutic drugs and the individuals who do not respond to therapeutic drugs).Typically, the feature data of the entire group (for example, case and control) is sorted in ascending order based on the value of a single feature.
[0187] In some embodiments, the disease is advanced adenoma (AA), colon cancer (CRC), colorectal cancer, or inflammatory bowel disease.
[0188] The term "input feature" or "feature" refers to a variable used by a model to predict the output classification (label) of a sample, e.g., a status, sequence content (e.g., mutation), proposed data collection operation, or proposed treatment. A value of the variable can be determined for a sample and used to determine the classification. Examples of input features for genetic data include aligned variables associated with the alignment of sequence data (e.g., sequence reads) to the genome, and non-aligned variables, e.g., variables associated with the sequence content of sequence reads, protein or autoantibody measurements, or average methylation levels in a genomic region.
[0189] The value of variable can be determined for sample and can be used to determine classification.Examples of input features of genetic data include aligned variables that are related to the alignment of sequence data (such as sequence reads) with genome, and non-aligned variables, such as the sequence content of sequence reads, the measured values of protein or autoantibody, or the variable related to the average methylation level in genome region.In various examples, genetic features such as V-plot measurements, FREE-C, cfDNA measurements of transcription start site, and DNA methylation level of cfDNA fragments are used as input features of machine learning methods and models.
[0190] In one example, the sequencing information includes information about multiple genetic features such as, but not limited to, transcription start sites, transcription factor binding sites, open and closed chromatin states, and nucleosome positions or occupancies.
[0191] B. Data Analysis In some embodiments, the present disclosure provides systems, methods, or kits with data analysis implemented in software applications, computing hardware, or both. In various embodiments, the analysis application or system includes at least a data reception module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In one embodiment, the data reception module can include a computer system that connects laboratory hardware or instrumentation to a computer system that processes laboratory data. In one embodiment, the data preprocessing module can include a hardware system or computer software that performs operations on data in preparation for analysis. Examples of operations that can be applied to data in the preprocessing module include affine transformations, noise removal operations, data cleaning, reformatting, or subsampling. The data analysis module can specialize in analyzing genomic data from one or more genomic materials, for example, by taking assembled genomic sequences and performing probabilistic and statistical analyses to identify abnormal patterns associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module may use analytical methods derived from, for example, statistics, mathematics, or biology to support an understanding of the association between the identified abnormal patterns and health status, functional status, prognosis, or risk. The data visualization module may use methods of mathematical modeling, computer graphics, or rendering to create visual representations of the data that can facilitate understanding or interpretation of the results.
[0192] In various embodiments, machine learning methods are applied to distinguish between samples in a population of samples, hi one embodiment, machine learning methods are applied to distinguish between healthy samples and aggressive adenoma samples.
[0193] In one embodiment, the one or more machine learning operations used to train the methylation-based prediction engine include a generalized linear model, a generalized additive model, a non-parametric regression operation, a random forest classifier, a spatial regression operation, a Bayesian regression model, a time series analysis, a Bayesian network, a Gaussian network, a decision tree learning operation, an artificial neural network, a recurrent neural network, a reinforcement learning operation, a linear / non-linear regression operation, a support vector machine, a clustering operation, and a genetic algorithm operation.
[0194] In various embodiments, the computational method is selected from logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-SNE, multilayer perceptron (MLP), network clustering, neuro-fuzzy, and artificial neural networks.
[0195] In some embodiments, the methods disclosed herein can include computational analysis of nucleic acid sequence data from samples from an individual or from multiple individuals. The analysis can be based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference to identify inferred variants from the sequence data. Non-limiting examples of analytical methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline mutations or somatic mutations. In some embodiments, variants can refer to known variants. Known variants can be scientifically confirmed or reported in the literature. In some embodiments, variants can refer to hypothetical variants associated with biological changes. The biological changes can be known or unknown. In some embodiments, hypothetical variants can be reported in the literature but not yet biologically confirmed.
[0196] Alternatively, putative variants are not reported in the literature but can be inferred based on the computer analyses disclosed herein. In some embodiments, germline variants can refer to nucleic acids that give rise to spontaneous or normal mutations.
[0197] Natural or normal variations may include, for example, skin color, hair color, and normal body weight. In some embodiments, somatic mutations may refer to nucleic acids that cause acquired or aberrant mutations. Acquired or aberrant mutations may include, for example, cancer, obesity, diseases, conditions, disorders, and disorders. In some embodiments, the analysis may include distinguishing between germline variants. Germline variants may include, for example, private variants and somatic mutations. In some embodiments, the identified variants may be used by clinicians or other health professionals to improve healthcare methodologies, diagnostic accuracy, and cost reduction.
[0198] Further provided herein are improved methods and computing systems or software media that can distinguish between nucleic acid sequence errors introduced by amplification and / or sequencing techniques, somatic mutations, and germline variants. The provided methods can include simultaneously calling and scoring variants from aligned sequencing data of all samples obtained from a patient.
[0199] Samples obtained from subjects other than patients can also be used. Other samples can also be collected from subjects that have been previously analyzed by sequencing assay or target sequencing assay (i.e., target resequencing assay). The methods, computing systems, or software media disclosed herein can improve the identification and accuracy of mutations or mutations (e.g., germline or somatic, including copy number variations, single nucleotide variations, indels, and gene fusions), and can reduce the detection limit by reducing the number of false positive and false negative identifications.
[0200] C. Classifier generation In one embodiment, the present system and method provides a classifier that is generated based on feature information derived from analyzing methylation sequences from a biological sample such as cfDNA. The classifier forms part of a prediction engine for distinguishing groups within a population based on methylation sequence features identified in a biological sample such as cfDNA.
[0201] In one embodiment, the classifier is created by formatting similar portions of methylation information into a unified format and a unified scale, storing the normalized methylation information in a column-oriented database, training a methylation prediction engine by applying one or more machine learning operations to the stored normalized methylation information, wherein the methylation prediction engine maps one or more feature combinations to a particular population, applying the methylation prediction engine to the accessed field information to identify methylation associated with a group, and classifying the individuals into a group.
[0202] Specificity can be defined as the probability of a negative test among people without the disease: it is equal to the number of people without the disease who test negative divided by the total number of individuals without the disease.
[0203] In various embodiments, the model, classifier, or predictive test has a specificity of at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.
[0204] Sensitivity can be defined as the probability of a positive test among people without the disease: it is equal to the number of individuals with the disease who test negative divided by the total number of individuals with the disease.
[0205] In various embodiments, the model, classifier, or predictive test has a sensitivity of at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.
[0206] In one embodiment, the group is selected from healthy (asymptomatic) inflammatory bowel disease, AA, or CRC.
[0207] D. Digital Processing Device In some embodiments, the subject matter described herein may include a digital processing device or use thereof. In some embodiments, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform the device's functions. In some embodiments, the digital processing device may include an operating system configured to execute executable instructions. In some embodiments, the digital processing device may be optionally connected to a computer network. In some embodiments, the digital processing device may be optionally connected to the Internet to connect to the World Wide Web. In some embodiments, the digital processing device may be optionally connected to a cloud computing infrastructure. In some embodiments, the digital processing device may be optionally connected to an intranet. In some embodiments, the digital processing device may be optionally connected to a data storage device.
[0208] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, handheld computers, Internet appliances, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, booklet, slate, and convertible configurations.
[0209] In some embodiments, a digital processing device may include an operating system configured to execute executable instructions. For example, an operating system may include software, including programs and data, that manages the device's hardware and provides services for the execution of applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Non-limiting examples of suitable personal computer operating systems include Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems, such as GNU / Linux®. In some embodiments, the operating system may be provided by cloud computing, where cloud computing resources may be provided by one or more service providers.
[0210] In some embodiments, the device may include a storage device and / or memory device. A storage device and / or memory device may be one or more physical devices used to temporarily or permanently store data or programs. In some embodiments, the device may be volatile memory, requiring power to maintain stored information. In some embodiments, the device may be nonvolatile memory, capable of retaining stored information when power is not applied to the digital processing device. In some embodiments, the nonvolatile memory may include flash memory. In some embodiments, the nonvolatile memory may include dynamic random access memory (DRAM). In some embodiments, the nonvolatile memory may include ferroelectric random access memory (FRAM). In some embodiments, the nonvolatile memory may include phase change random access memory (PRAM). In some embodiments, the device may be a storage device, including, for example, a CD-ROM, a DVD, a flash memory device, a magnetic disk drive, a magnetic tape drive, an optical disk drive, and a cloud computing-based storage device. In some embodiments, the storage device and / or memory device may be a combination of devices, such as those disclosed herein.
[0211] In some examples, the digital processing device may include a display for conveying visual information to a user. In some embodiments, the display may be a cathode ray tube (CRT). In some embodiments, the display may be a liquid crystal display (LCD). In some embodiments, the display may be a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display may be an organic light emitting diode (OLED) display. In some embodiments, the OLED display may be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some embodiments, the display may be a plasma display. In some embodiments, the display may be a video projector. In some embodiments, the display may be a combination of devices such as those disclosed herein.
[0212] In some embodiments, the digital processing device may include an input device for receiving information from a user. In some embodiments, the input device may be a keyboard. In some embodiments, the input device may be a pointing device, including, for example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some embodiments, the input device may be a touchscreen or multi-touchscreen. In some embodiments, the input device may be a microphone that captures voice or other audio input. In some embodiments, the input device may be a video camera that captures motion or visual input. In some embodiments, the input device may be a combination of devices, such as those disclosed herein.
[0213] E. Non-Transitory Computer-Readable Storage Medium In some embodiments, the subject matter disclosed herein may include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of an optionally networked digital processing device. In some embodiments, the computer-readable storage medium may be a tangible component of the digital processing device. In some embodiments, the computer-readable storage medium may be optionally removable from the digital processing device. In some embodiments, the computer-readable storage medium may include, for example, a CD-ROM, a DVD, a flash memory device, a solid-state memory, a magnetic disk device, a magnetic tape drive, an optical disk drive, a cloud computing system and service, or the like. In some embodiments, the program and instructions may be encoded on the medium permanently, nearly permanently, semi-permanently, or non-transitoryly.
[0214] F. Computer Systems The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 5 shows a computer system (501) programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, or reference sequences. The computer system (501) can process various aspects of the patient data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system (501) can be a user's electronic device or a computer system located remotely relative to the electronic device. The electronic device can be a mobile electronic device.
[0215] The computer system (501) includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") (505), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (501) also includes memory or storage locations (510) (e.g., random access memory, read-only memory, flash memory), electronic storage devices (515) (e.g., hard disks), communication interfaces (520) (e.g., network adapters) for communicating with one or more other systems, and peripheral devices (525), such as cache, other memory, data storage devices, and / or electronic display adapters. The memory (510), storage devices (515), interfaces (520), and peripheral devices (525) communicate with the CPU (505) via a communication bus (solid lines), such as a motherboard. The storage devices (515) may be a data storage device (or data repository) for storing data. The computer system (501) may be operatively connected to a computer network ("network") (530) with the aid of a communication interface (520). The network (530) may be the Internet and / or an extranet, or an intranet and / or extranet in communication with the Internet. In some embodiments, the network (530) is a telecommunications and / or data network. The network (530) may include one or more computer servers, which may enable distributed computing, such as cloud computing. The network (530), in some embodiments, may implement a peer-to-peer network with the aid of the computer system (501), thereby enabling devices coupled to the computer system (501) to act as clients or servers.
[0216] The CPU (505) can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a storage location, such as memory (510). The instructions may be directed to the CPU (505), which can then program or otherwise configure the CPU (505) to implement the methods of the present disclosure. Examples of operations performed by the CPU (505) include fetch, decode, execute, and writeback.
[0217] The CPU 505 may be part of a circuit, such as an integrated circuit. One or more other components of the system 501 may also be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).
[0218] The storage device (515) can store files such as drivers, libraries, and saved programs. The storage device (515) can store user data, such as user preferences and user programs. The computer system (501) can, in some embodiments, include one or more additional data storage devices external to the computer system (501), such as located on a remote server in communication with the computer system (501) via an intranet or the Internet.
[0219] The computer system (501) can communicate with one or more remote computer systems via the network (530). For example, the computer (501) can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system (501) via the network (530).
[0220] The methods described herein can be performed by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system (501), such as, for example, on memory (510) or electronic storage (515). The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor (505). In some embodiments, the code can be retrieved from storage (515) and stored in memory (510) for immediate access by the processor (505). In some embodiments, the electronic storage (515) can be eliminated, and machine-executable instructions are stored in memory (510).
[0221] The code may be pre-compiled and configured for use with a machine having a suitable processor to execute the code, or may be interpreted or compiled at run time. The code may be supplied in a programming language that may be selected to render the code executable in a pre-compiled, interpreted, or as-compiled fashion.
[0222] Aspects of the systems and methods provided herein, such as the computer system (501), may be embodied in programming. Various aspects of the technology may be thought of as "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data executed or embodied on a type of machine-readable medium. The machine-executable code may be stored in electronic storage devices such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory recording media at any time for programming the software. All or portions of the software are sometimes communicated via the Internet or various other telecommunications networks. Such communication may enable, for example, loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and on various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to media that participate in providing instructions to a processor for execution.
[0223] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, any of the storage devices in a computer or the like, such as those that may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, DVDs, or DVD-ROMs, other optical media, punch cards, paper tape, other physical storage media with patterns of holes, RAM, ROM, PROMs, and EPROMs, FLASH®-EPROMs, other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0224] The computer system (501) may include or be in communication with an electronic display (135) that includes a user interface (UI) (540) for providing, for example, nucleic acid sequences, enriched nucleic acid samples, expression profiles, and analysis of the expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0225] The disclosed methods and systems can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit (505). The algorithms can, for example, search for multiple regulatory elements, sequence a nucleic acid sample, enrich a nucleic acid sample, determine an expression profile of a nucleic acid sample, analyze the expression profile of a nucleic acid sample, and record or disseminate results of the analysis of the expression profile.
[0226] In some embodiments, the subject matter disclosed herein includes at least one computer program or the use of such a computer program. A computer program may be a set of instructions that can be executed by a CPU, GPU, or TPU of a digital processing device and written to perform a specific task. The computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), data structures, etc., that perform a specific task or extract a specific data type. Given the disclosure provided herein, those skilled in the art will recognize that computer programs can be written in a variety of languages and in various versions.
[0227] The functionality of the computer-readable instructions may be combined or distributed as needed for various environments. In some embodiments, a computer program may include one sequence of instructions. In some embodiments, a computer program may include multiple sequences of instructions. In some embodiments, a computer program may be provided from one location. In some embodiments, a computer program may be provided from multiple locations. In some embodiments, a computer program may include one or more software modules. In some embodiments, a computer program may include, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or any combination thereof.
[0228] In some embodiments, the computational processing may be a method of statistics, mathematics, biology, or any combination thereof. In some embodiments, the computational processing method comprises a dimension reduction method, including, for example, logistic regression, dimensionality reduction, principal component analysis, autoencoder, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosted trees, logistic regression, matrix factorization, network clustering, neural networks.
[0229] In some embodiments, the computational methods are supervised machine learning methods, including, for example, regression, support vector machines, tree-based methods, and networks.
[0230] In some embodiments, the computational methods are unsupervised machine learning methods, including, for example, clustering, networks, principal component analysis, and matrix factorization.
[0231] G. Database In some embodiments, the subject matter disclosed herein includes one or more databases, or the use thereof, for storing patient data, biological data, biological sequences, or reference sequences. The reference sequences can be derived from a database. Given the disclosure provided herein, one of skill in the art will recognize that many databases are suitable for storing and retrieving sequence information. In some embodiments, suitable databases can include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some embodiments, the database can be internet-based. In some embodiments, the database can be web-based. In some embodiments, the database can be cloud computing-based. In some embodiments, the database can be based on one or more local computer storage devices.
[0232] V. Cancer Diagnosis and Detection The trained machine learning methods, models, and discriminative classifiers described herein are useful for a variety of medical applications, including cancer detection, diagnosis, and treatment response. Once the models are trained using individual metadata and analysis-derived features, the applications can be adapted to stratify individuals in a population and guide treatment decisions accordingly.
[0233] A. Diagnosis The methods and systems provided herein can perform predictive analysis using an artificial intelligence-based approach to analyze data obtained from a subject (patient) to generate an output of a diagnosis for the subject with cancer (e.g., CRC). For example, the application can apply a predictive algorithm to the obtained data to generate a diagnosis for the subject with cancer. The predictive algorithm can include an artificial intelligence-based predictor, such as a machine learning-based predictor, configured to process the obtained data to generate a diagnosis for the subject with cancer.
[0234] The machine learning predictor can be trained using datasets from one or more sets of cohorts of cancer patients as input to the machine learning predictor and known diagnostic (e.g., staging and / or tumor proportion) outcomes for the subjects as output, e.g., datasets generated by performing a multi-analytical assay on individual biological samples.
[0235] A training dataset (e.g., a dataset generated by performing a multi-analysis assay on an individual's biological samples) can be generated, for example, from one or more sets of subjects with common characteristics (features) and results (labels). The training dataset can include a set of features and labels corresponding to diagnostically relevant features. Features can include, for example, a range or category of cfDNA assay measurements, such as the number of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap or fall within each of a set of bins (genomic windows) of a reference genome. For example, a set of features collected from a given subject at a given time point can collectively function as a diagnostic signature and indicate the subject's identified cancer at a given time point. Features can also include labels indicative of a subject's diagnostic outcome, such as for one or more cancers.
[0236] The label can include, for example, an outcome, such as a known diagnosis (e.g., stage classification and / or tumor proportion) of the subject. The outcome can include a characteristic associated with cancer in the subject. For example, the characteristic can indicate that the subject suffers from one or more cancers.
[0237] The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more sets of subjects (e.g., a retrospective and / or prospective cohort of patients with or without one or more cancers). Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more sets of subjects (e.g., a retrospective and / or prospective cohort of patients with or without one or more cancers). The training set may be balanced across multiple sets of data corresponding to one or more sets of subjects (e.g., patients from various clinical centers or clinical trials). The machine learning predictor may be trained until a predefined condition for accuracy or performance is met, such as having a minimum target value corresponding to a diagnostic accuracy measure. For example, the diagnostic accuracy measure may correspond to the diagnosis, staging, or prediction of tumor fraction of one or more cancers in a subject.
[0238] Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC) of the ROC curve, which correspond to diagnostic accuracy in detecting or predicting cancer (e.g., colon cancer).
[0239] In another aspect, the present disclosure provides a method for identifying cancer in a subject, the method comprising: (a) providing a biological sample comprising cell-free nucleic acid (cfNA) molecules from the subject; (b) methylation sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating quantitative measures of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 distinct regions (each of the at least about 10 distinct regions); and (e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has cancer.
[0240] For example, such predefined conditions may include a sensitivity for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0241] As another example, such a predefined condition may be a specificity for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0242] As another example, such a predefined condition may be a positive predictive value (PPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0243] As another example, such a predefined condition may include a negative predictive value (NPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0244] As another example, such a predefined condition may be that the AUC of a ROC curve predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0245] In some examples of any of the aforementioned aspects, the method further includes monitoring the progression of a disease in the subject, wherein said monitoring is based at least in part on the gene sequence characteristics. In some examples, the disease is cancer.
[0246] In some examples of any of the aforementioned aspects, the method further includes determining a tissue origin of the cancer in the subject, wherein said determining is based at least in part on the gene sequence characteristics.
[0247] In some examples of any of the aforementioned aspects, the method further includes estimating tumor burden in the subject, wherein said estimating is based at least in part on the gene sequence characteristics.
[0248] B. Treatment Response The predictive classifiers, systems, and methods described herein are useful for classifying populations of individuals for a number of clinical applications (e.g., based on performing multi-analytical assays on biological samples of the individuals), including detecting early cancer, diagnosing cancer, classifying cancer into specific stages of the disease, or determining responsiveness or resistance to therapeutic agents for treating cancer.
[0249] The methods and systems described herein are applicable to various cancer types, as well as grades and stages, and are therefore not limited to a single cancer disease type. Thus, a combination of analyses and assays can be used in the present systems and methods to predict response to cancer therapy across various cancer types in various tissues and to classify individuals based on treatment response. In one example, the classifiers described herein stratify groups of individuals into treatment responders and non-responders.
[0250] The present disclosure further provides a method for determining drug targets (e.g., genes associated / significant to a particular class) for a disease or disorder of interest, the method comprising: evaluating a sample obtained from an individual for gene expression levels of at least one gene; and using a proximity analysis routine to determine genes associated with the classification of the sample, thereby ascertaining one or more drug targets associated with the classification.
[0251] The present disclosure further provides a method for determining the effectiveness of a drug designed to treat a disease class, the method comprising: obtaining a sample from an individual having the disease class; exposing the sample to the drug; evaluating the drug-exposed sample for gene expression levels of at least one gene; and using a computer model constructed using a weighted voting scheme to classify the drug-exposed sample into the disease class according to the relative gene expression levels of the sample to the relative gene expression levels of a model.
[0252] The present disclosure further provides a method for determining the effectiveness of a drug designed to treat a disease class, wherein an individual has been exposed to the drug, the method comprising: obtaining a sample from the individual exposed to the drug; evaluating the sample for a gene expression level of at least one gene; and using a model constructed using a weighted voting scheme to classify the sample into the disease class, wherein the using comprises evaluating the gene expression levels of the sample compared to the gene expression levels of a model.
[0253] However, another application is a method for determining whether an individual belongs to a phenotypic class (e.g., intelligence, response to treatment, longevity, susceptibility to viral infection, or obesity), the method comprising obtaining a sample from the individual, assessing the sample for gene expression levels of at least one gene, and using a model constructed using a weighted voting scheme to classify the sample into a disease class, comprising assessing the gene expression levels of the sample compared to the gene expression levels of a model.
[0254] Biomarkers can be useful in predicting the prognosis of colorectal cancer patients. The ability to classify patients as high-risk (poor prognosis) or low-risk (good prognosis) makes it possible to select appropriate treatment for these patients. For example, high-risk patients may benefit from aggressive treatment, while low-risk patients may not see significant benefit from treatment.
[0255] There are predictive biomarkers that can guide treatment decisions by identifying subsets of patients who are likely to be "exceptional responders" to a particular cancer therapy, or who may benefit from alternative treatment modalities.
[0256] In one aspect, the system and method described herein for classifying populations based on treatment response refers to, but is not limited to, cancers treated by chemotherapy agents of class DNA damaging agents, DNA repair target therapy, inhibitors of DNA damage signal transduction, inhibitors of DNA damage-induced cell cycle arrest, and the inhibition of processes that are indirectly linked to DNA damage.Each of these chemotherapy agents can be considered as " DNA damaging therapeutic agents ".
[0257] The patient's analyte data can be classified into high-risk and low-risk patient groups, such as patients at high or low risk of clinical recurrence, and the results can be used to determine treatment strategies. For example, patients determined to be high-risk patients may be treated with adjuvant chemotherapy after surgery. For patients considered to be low-risk patients, adjuvant chemotherapy may be withheld after surgery. Thus, in some aspects, the present disclosure provides a method for preparing a gene expression profile of a colorectal cancer tumor indicative of recurrence risk.
[0258] In various examples, the classifiers described herein stratify a population of individuals between responders and non-responders to a treatment.
[0259] In various examples, the treatment is selected from alkylating agents, plant alkaloids, antitumor antibiotics, antimetabolites, topoisomerase inhibitors, retinoids, checkpoint inhibitor therapy, and VEGF inhibitors.
[0260] Examples of treatments for which populations may be stratified into responders and non-responders include, but are not limited to, sorafenib, regorafenib, imatinib, eribulin, gemcitabine, capecitabine, pazopanib, lapatinib, dabrafenib, sunitinib malate, crizotinib, everolimus, torisirolimus, sirolimus, axitinib, gefitinib, anastrozole, bicalutamide, fulvestrant, ralitrexed, pemetrexed, goserelin acetate, erlotinib, vemurafenib, bisphosphonate ... Modegib, tamoxifen citrate, paclitaxel, docetaxel, cabazitaxel, oxaliplatin, ziv-aflibercept, bevacizumab, trastuzumab, pertuzumab, panitumumab, taxanes, bleomycin, melphalen, plumbagin, camptosar, mitomycin C, mitoxantrone, poly(styrene maleic acid)-bound neocarzinostatin (SMANCS), doxorubicin, pegylated doxorubicin, FOLFORI, 5-fluorouracil temozolomide tegafur, gimeracil, oteraci, itraconazole, bortezomib, lenalidomide, irinotecan, epirubicin, romidepsin, resminostat, tasquinimod, refametinib, lapatinib, Tyverb®, Arenegyr, NGR-TNF, pasireotide, Signifor®, ticilimumab, tremelimumab, lansoprazole, PrevOnco®, ABT-869, linifanib, vorolanib, tivantinib, tal These include chemotherapy agents including Ceba®, erlotinib, Stivarga®, regorafenib, fluoro-sorafenib, brivanib, liposomal doxorubicin, lenvatinib, ramucirumab, peretinoin, Ruchiko, muparfostat, Teysuno®, tegafur, gimeracil, oteracil, and orantinib, and antibody therapies including alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, or rituximab.
[0261] In other examples, a population may be stratified into responders and non-responders to checkpoint inhibitor treatment, such as compounds that bind to PD-1 or CTLA4.
[0262] In another example, a population may be stratified into responders and non-responders to anti-VEGF therapy that binds to a VEGF pathway target.
[0263] VI. Indications In some examples, the biological state may include a disease. In some examples, the biological state may be a stage of a disease. In some examples, the biological state may be a gradual change in a biological state. In some examples, the biological state may be a treatment effect. In some examples, the biological state may be a drug effect. In some examples, the biological state may be a surgical effect. In some examples, the biological state may be a biological state after a lifestyle modification. Non-limiting examples of lifestyle modifications include a change in diet, a change in smoking, and a change in sleep patterns. In some examples, the biological state is unknown. The analyses described herein may include machine learning to infer an unknown biological state or to interpret an unknown biological state.
[0264] In one example, the present systems and methods are particularly useful for applications related to colon cancer, i.e., cancer that originates in the tissues of the colon (the longest part of the large intestine). Most colon cancers are adenocarcinomas (cancers that begin in cells that form linear organs and have glandular properties). Cancer progression is characterized by the stage or extent of cancer in the body. Staging is usually based on the size of the tumor, whether lymph nodes contain cancer, and whether the cancer has spread from the site of initial development to other parts of the body. Stages of colon cancer include stage I, stage II, stage III, and stage IV. Unless otherwise specified, the term "colon cancer" refers to stage 0, stage I, stage II (including stage IIA or IIB), stage III (including stage IIIA, IIIB, or IIIC), or stage IV colon cancer. In some examples described herein, the colon cancer is of any stage. In some examples, the colon cancer is stage I colon cancer. In some examples, the colon cancer is stage II colon cancer. In some instances, the colon cancer is stage III colon cancer. In some instances, the colon cancer is stage IV colon cancer.
[0265] Diseases that can be inferred by the disclosed methods include, for example, cancer, gut-related diseases, immune-mediated inflammatory diseases, nervous system diseases, kidney diseases, prenatal diseases, and metabolic diseases.
[0266] In some examples, the methods of the present disclosure can be used to diagnose cancer, non-limiting examples of which include adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0267] Non-limiting examples of cancers that can be inferred by the disclosed methods and systems include acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brain stem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumor, breast cancer, bronchial tumor, Burkitt's lymphoma, non-Hodgkin's lymphoma, carcinoid tumor, cervical cancer, chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing's sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrous histiocytoma, gallbladder cancer, Gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin's lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell lung cancer, melanoma, oral cancer, myelodysplastic syndrome, multiple myeloma, medulloblastoma, nasal cavity cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papilloma, paraganglioma, These include parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumors, plasma cell neoplasms, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome, skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, and Wilms' tumor.
[0268] Non-limiting examples of bowel-related diseases that can be inferred by the disclosed methods and systems include Crohn's disease, colitis, ulcerative colitis (UC), inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), and celiac disease. In some examples, the disease is inflammatory bowel disease, colitis, ulcerative colitis, Crohn's disease, microscopic colitis, collagenous colitis, lymphocytic colitis, diversion colitis, Behcet's disease, and ulcerative colitis.
[0269] Non-limiting examples of immune-mediated inflammatory diseases that can be inferred by the disclosed methods and systems include psoriasis, sarcoidosis, rheumatoid arthritis, asthma, rhinitis (hay fever), food allergies, eczema, lupus, multiple sclerosis, fibromyalgia, type 1 diabetes, and Lyme disease. Non-limiting examples of nervous system diseases that can be inferred by the disclosed methods and systems include Parkinson's disease, Huntington's disease, multiple sclerosis, Alzheimer's disease, stroke, epilepsy, neurodegeneration, and neuropathy. Non-limiting examples of kidney diseases that can be inferred by the disclosed methods and systems include interstitial nephritis, acute renal failure, and nephropathy. Non-limiting examples of prenatal diseases that can be inferred by the disclosed methods and systems include Down syndrome, aneuploidy, spina bifida, trisomy, Edwards syndrome, teratoma, sacrococcygeal teratoma (SCT), ventriculomegaly, renal agenesis, cystic fibrosis, and hydrops fetalis. Non-limiting examples of metabolic diseases that can be inferred by the disclosed methods and systems include cystinosis, Fabry disease, Gaucher disease, Lesch-Nyhan syndrome, Niemann-Pick disease, phenylketonuria, Pompe disease, and Tay-Sachs disease.
[0270] The specific details of particular examples may be combined in any suitable manner without departing from the spirit and scope of the disclosed examples of the invention. However, other examples of the invention may be directed to particular examples of individual aspects or particular combinations of these individual aspects. All patents, patent applications, publications, and descriptions referred to herein are incorporated by reference in their entirety for all purposes.
[0271] VII. Kit The present disclosure provides a kit for identifying or monitoring cancer in a subject. The kit includes probes for identifying quantitative measures (e.g., indicating presence, absence, or relative abundance) of sequences at each of a plurality of cancer-associated genomic loci in a cell-free biological sample from the subject. The quantitative measures (e.g., indicating presence, absence, or relative abundance) of sequences at each of a plurality of cancer-associated genomic loci in the cell-free biological sample can be indicative of one or more cancers. The probes can be selective for sequences at a plurality of cancer-associated genomic loci in the cell-free biological sample. The kit includes instructions for processing the cell-free biological sample using the probes to generate a dataset indicating quantitative measures (e.g., indicating presence, absence, or relative abundance) of sequences at each of a plurality of cancer-associated genomic loci in the cell-free biological sample from the subject. In one embodiment, the kit includes a primer set, PCR reaction components, sequencing reagents, minimally disruptive conversion reagents, and library preparation reagents.
[0272] The probes in the kit may be selective for sequences at multiple cancer-associated genomic loci in a cell-free biological sample. The probes in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with nucleic acid sequences from one or more of the multiple cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more different cancer-associated genomic loci or genomic regions identified for targeted methylation sequencing.
[0273] The instructions in the kit include instructions for analyzing the acellular biological sample using probes selective for sequences at multiple cancer-associated genomic loci in the acellular biological sample. These probes can be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) from one or more of the multiple cancer-associated genomic loci. These nucleic acid molecules can be primers or enrichment sequences. The instructions for analyzing the acellular biological sample can include instructions for performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the acellular biological sample to generate a dataset that indicates a quantitative measure (e.g., indicating the presence, absence, or relative amount) of the sequence at each of the multiple cancer-associated genomic loci in the acellular biological sample. The quantitative measure (e.g., indicating the presence, absence, or relative amount) of the sequence at each of the multiple cancer-associated genomic loci in the acellular biological sample can indicate one or more cancers.
[0274] The kit includes instructions for measuring and interpreting assay readouts, which may be quantified at one or more of a plurality of cancer-associated genomic loci to generate a dataset indicating a quantitative measure (e.g., indicating the presence, absence, or relative amount) of a sequence at each of a plurality of cancer-associated genomic loci in a cell-free biological sample. For example, array hybridization or polymerase chain reaction (PCR) corresponding to a plurality of cancer-associated genomic loci may be quantified to generate a dataset indicating a quantitative measure (e.g., indicating the presence, absence, or relative amount) of a sequence at each of a plurality of cancer-associated genomic loci in a cell-free biological sample. Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized versions thereof. [Example]
[0275] Example 1: Preparation of targeted EM-seq libraries and generation of classifiers.
[0276] Starting material: 10–200 ng double-stranded DNA.
[0277] 1.DNA preparation EDTA was removed from the DNA before oxidation, and the DNA samples had a final volume of 29 μl. Control DNA was used to assess oxidation and deamination. For sequencing on the Illumina platform, the Enzymatic Methyl-seq Kit Manual (NEB #E7120) was consulted for usage recommendations.
[0278] 2. Adapter Ligation 3. Oxidation of 5-methylcytosine and 5-hydroxymethylcytosine TET2 buffer was prepared. The TET2 reaction mixture was then added to one tube containing the TET2 Reaction Buffer Supplement and mixed thoroughly. On ice, the TET2 Reaction Buffer, oxidation supplement, pro-oxidant, and TET2 enzyme were added directly to the DNA sample. The mixture was then thoroughly mixed by vortexing. After a brief centrifugation, diluted Fe(II) solution was added to the mixture. The mixture was then thoroughly mixed by vortexing or pipetting up and down and briefly centrifuged. The mixture was then incubated at 37°C for 1 hour in a thermocycler. The sample was then transferred to ice and treated with 1 μl of stop reagent (yellow). The mixture was then thoroughly mixed by vortexing or pipetting up and down at least 10 times and briefly centrifuged. Finally, the mixture was incubated at 37°C for 30 minutes in a thermocycler, followed by 4°C.
[0279] 4. Cleanup of TET2-converted DNA The sample purification beads were resuspended by vortexing. Next, NEBNext sample purification beads were added to each sample and mixed thoroughly by pipetting up and down. The samples were incubated on the benchtop at room temperature for at least 5 minutes. The tubes were then placed on an appropriate magnetic stand to separate the beads from the supernatant. After 5 minutes (or when the solution became clear), the supernatant was carefully removed and discarded without disturbing the beads containing the DNA targets. While still on the magnetic stand, freshly prepared 80% ethanol was added to each of the tubes. The samples were incubated at room temperature for 30 seconds, and then the supernatant was carefully removed and discarded. This wash was repeated once for a total of two washes. After the second wash, all visible liquid was removed using a p10 pipette tip. The beads were then allowed to air dry for 2 minutes while the tubes were still on the magnetic stand with the lids open. The tubes were then removed from the magnetic stand. DNA was eluted from the beads using elution buffer. Elution buffer was added to each tube and mixed thoroughly by pipetting up and down 10 times. The samples were then incubated at room temperature for at least 1 minute. If necessary, the samples were briefly centrifuged to collect liquid from the sides of the tube before returning the tube to the magnetic stand. The tubes were then returned to the magnetic stand. After 3 minutes (or whenever the solution was clear), the eluted DNA from the supernatant was transferred to a new PCR tube.
[0280] 5. DNA Denaturation Prior to cytosine deamination, DNA was denatured using either formamide or 0.1 N sodium hydroxide.
[0281] 6. Cytosine deamination On ice, APOBEC reaction buffer, BSA, and APOBEC were added to the denatured DNA. The mixture was then thoroughly mixed by vortexing or pipetting up and down at least 10 times before being briefly centrifuged. The mixture was then incubated in a thermocycler at 37°C for 3 hours, followed by incubation at 4°C.
[0282] 7. Cleanup of Deaminated DNA The sample purification beads were resuspended by vortexing. Next, 100 μl of resuspended NEBNext sample purification beads was added to each sample and then thoroughly mixed by pipetting up and down at least 10 times. During the final mix, all liquid was carefully expelled from the tip. The samples were then incubated on the benchtop at room temperature for at least 5 minutes. After 5 minutes (or when the solution became clear), the supernatant was carefully removed and discarded. While still on the magnetic stand, freshly prepared 80% ethanol was added to the tubes. The samples were then incubated at room temperature for 30 seconds, after which the supernatant was carefully removed and discarded. This wash was repeated once for a total of two washes. Next, the beads were air-dried for 90 seconds with the tubes still on the magnetic stand with the lids open. The DNA target was then eluted from the beads using elution buffer. Elution buffer was added to each tube and thoroughly mixed by pipetting up and down 10 times. The samples were incubated at room temperature for at least 1 minute. If necessary, the sample was briefly centrifuged to collect liquid from the sides of the tube before returning the tube to the magnetic stand. After 3 minutes (or whenever the solution became clear), the eluted DNA target in the supernatant was transferred to a new PCR tube.
[0283] 8. Multiplex Amplification and Targeted Methylation Classification To enable targeted methylation analysis of pre-identified regions of the genome, raw data files were used for alignment and methylation calling using conventional tools. Whole-genome amplification of enzymatically converted DNA was performed. Target enrichment was performed on the enzymatically converted library, and pre-identified DNA fragments containing target CpG sites were specifically pulled down using 5'-biotinylated capture probes. Hybrid selection was performed using the Illumina TruSightVR Rapid Capture Kit. Capture Target Buffer 3 (Illumina) was used instead of enrichment hybridization buffer during the hybridization step. After hybridization, the captured DNA fragments were amplified with 14 PCR cycles. The target capture library was sequenced on an Illumina HiSeqVR 2500 sequencer using 2x100-cycle runs with 4-5 samples in rapid run mode. Spike-in of 10% PhiX into the enzymatic sequencing library increased base diversity and improved sequencing quality.
[0284] Using conventional methods, FASTQ files were mapped to a reference genome, and methylation scores were calculated for disease classification. Featurized data containing a set of CpG sites associated with health, disease, disease state, and treatment response were input into a machine learning model to identify classifiers that stratify individuals within the population.
[0285] Example 2: Targeted EM-seq using a conversion-resistant sequencing adapter / primer system Identified adapters of known sequences were ligated to the ends of DNA molecules in samples with unknown sequences. The adapters were then used to PCR amplify the entire library of diverse molecules using a single set of primers corresponding to the known adapters. The ligated adapter sequences were further used as binding sites for sequencing primers during subsequent sequencing reactions. To utilize the data provided by duplex sequencing, partially double-stranded adapters with unique molecular identifiers (UMIs) were ligated to double-stranded DNA.
[0286] To improve the robustness of EM-Seq duplex sequencing (and reduce costs) against oxidation efficiency, conversion-resistant adapters can be used to increase the consistency of sequencing library quality. Conversion-resistant adapters contain only unmodified bases, allowing full base conversion of the adapter. An example of a conversion-resistant adapter is shown in Panel A of Figure 4.
[0287] In the absence of conversion, sequencing libraries generated using these conversion-resistant adapters can be amplified and sequenced using a pair of PCR and sequencing primers that match the original adapter sequences. In the presence of conversion, sequencing libraries can be amplified and sequenced using PCR and sequencing primers that match the converted adapter sequences, as shown in Figure 4, panel B.
[0288] One functional example set of conversion-resistant adapters, PCR primers, and sequencing primers was tested (Figure 6). Sequencing library yields for libraries generated using either conversion-resistant adapters or 5mC-containing adapters are shown. Because the TET-mediated oxidation step was not performed, all Cs and 5mCs were susceptible to C-to-U conversion. While the 5mC-containing adapter system was more efficient without conversion, the conversion-resistant adapter required conversion so that the conversion-resistant adapter system could be amplified using conversion-specific PCR primers. The DNA sequences for these conversion-specific PCR primers are listed in Table 1.
[0289] [Table 1-1]
[0290] [Table 1-2]
[0291] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided within the specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It will be understood that various alternatives to the embodiments of the present invention described herein may be utilized in practicing the invention. It is therefore contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.
Claims
1. 1. A method for performing methylation sequencing on nucleic acid molecules of a biological sample, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to the nucleic acid molecule, wherein the nucleic acid molecule comprises unconverted nucleic acid; b) converting unmethylated cytosines to uracils in said nucleic acid molecule using a minimally disruptive conversion method, thereby producing a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction, thereby producing an amplified converted nucleic acid; d) probing the amplified converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby producing probed converted nucleic acid; e) determining the nucleic acid sequence of the probed converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequences of the probed converted nucleic acids with reference nucleic acid sequences of a pre-identified panel of CpG or CH loci to determine the methylation profile of the nucleic acid molecules of the biological sample.
2. 2. The method of claim 1, wherein the nucleic acid molecule is plasma cfDNA.
3. 10. The method of claim 1, wherein the minimally disruptive conversion method comprises enzymatic conversion, TAPS, or CAPS.
4. 2. The method of claim 1, wherein the unique molecular identifier is 4 bp to 6 bp in length and has a 5' thymidine overhang.
5. 5. The method of claim 4, wherein the nucleic acid adapter further comprises a unique dual index (UDI) sequence.
6. 6. The method of claim 5, wherein the UDI sequence has a length of 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp.
7. 2. The method of claim 1, wherein amplifying the converted nucleic acid comprises using a primer comprising a unique dual index (UDI) sequence.
8. 2. The method of claim 1, wherein the nucleic acid adapter is a conversion-resistant adapter that contains guanine, thymine, adenine, and cytosine bases but does not contain 5mC- or 5hmC-containing bases.
9. The method of claim 1 , wherein the nucleic acid probe is an unmethylated nucleic acid probe.
10. 10. The method of claim 1, wherein the nucleic acid probe hybridizes to a target region of interest that matches an unmethylated cytosine at a CpG site in a reference nucleic acid sequence.
11. 2. The method of claim 1, wherein the nucleic acid probe comprises a target region of interest that matches a methylated cytosine at a CpG site in a reference nucleic acid sequence.
12. 2. The method of claim 1, wherein the nucleic acid probe is a mixture of chemically or enzymatically modified methylated or unmethylated nucleic acid probes.
13. 2. The method of claim 1, wherein one or more cytosines in the CG context of the probed converted nucleic acid are converted to thymine, and all cytosines in the CH context of the probed converted nucleic acid are converted to thymine.
14. 2. The method of claim 1, wherein the conversion of unmethylated cytosine to uracil comprises a series of TET / APOBEC enzymatic conversions.
15. 2. The method of claim 1, wherein the conversion of unmethylated cytosine to uracil comprises TAPS.
16. 1. A method for determining targeted methylation patterns within nucleic acid molecules of a biological sample from a subject, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to the nucleic acid molecule, wherein the nucleic acid molecule comprises unconverted nucleic acid; b) enzymatically converting unmethylated cytosines to uracils within the nucleic acid molecule to produce converted nucleic acids; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to said pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of the converted nucleic acid with reference nucleic acid sequences of a pre-identified panel of CpG or CH loci to determine a targeted methylation pattern of nucleic acid molecules in a biological sample from the subject.
17. 17. The method of claim 16, wherein determining the nucleic acid sequence of the converted nucleic acid comprises double-stranded error correction.
18. 17. The method of claim 16, wherein the nucleic acid adapter is a conversion-resistant adapter that includes guanine, thymine, adenine, and cytosine bases but does not include 5mC- or 5hmC-containing bases.
19. 17. The method of claim 16, wherein the pre-identified panel of CpG or CH loci comprises loci associated with transcription factor start sites.
20. 17. The method of claim 16, wherein the targeted methylation pattern comprises hemimethylated CpG loci.
21. 1. A method for determining a methylation profile of a cell-free DNA (cfDNA) sample from a subject, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to cfDNA, wherein the cfDNA comprises unconverted nucleic acid; b) enzymatically converting unmethylated cytosines to uracils within the nucleic acid molecule to produce converted nucleic acids; c) amplifying the converted nucleic acid by polymerase chain reaction; d) probing the converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to said pre-identified panel of CpG or CH loci; e) determining the nucleic acid sequence of the converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequence of said converted nucleic acid to reference nucleic acid sequences of said pre-identified panel of CpG or CH loci to determine a methylation profile of a cell-free DNA (cfDNA) sample from the subject.
22. 22. The method of claim 21, wherein the nucleic acid adapter is a conversion resistant adapter that includes guanine, thymine, adenine, and cytosine bases but does not include 5mC- or 5hmC-containing bases.
23. 22. The method of claim 21, further comprising identifying a cfDNA sample from a tissue, identifying somatic mutations in the cfDNA sample, estimating nucleosome positions in the cfDNA sample, identifying variably methylated regions in the cfDNA sample, or identifying haplotype blocks in the cfDNA sample.
24. 1. A method for performing methylation sequencing on nucleic acid molecules of a biological sample, the method comprising: a) preparing a methylation sequencing library from cfDNA fragments of a nucleic acid molecule, i) ligating double-stranded adapters to the cfDNA fragments; ii) ligating a dual unique molecular identifier to the cfDNA fragment; and iii) preparing a methylation sequencing library from the cfDNA nucleic acid molecule by converting unmethylated cytosines in the cfDNA fragments to uracils using a minimally disruptive conversion method; b) enriching the methylation sequencing library for sequences corresponding to CpG or CH loci, thereby generating an enriched methylation sequencing library; c) sequencing the enriched methylation sequencing library at a depth of greater than 100x using single-end or paired-end reads, thereby generating single-end or paired-end read sequencing fragments; d) for each sequencing fragment of the paired-end reads, correcting sequencing errors within the overlapping region of the paired-end reads; e) collapsing the sequencing fragments into concatenated read families to correct errors resulting from PCR and sequencing; f) collapsing said concatenated read families into dual read families to identify methylation discrepancies relative to the predicted methylation status of symmetric CpG loci in said nucleic acid molecule.
25. 25. The method of claim 24, wherein the minimally disruptive conversion method is enzymatic conversion, TAPS, or CAPS.
26. 1. A method for generating a classifier, the method comprising: a) ligating a nucleic acid adapter comprising a unique molecular identifier to each nucleic acid molecule of a biological sample from a healthy subject and a biological sample from a cancer subject; b) converting unmethylated cytosines to uracils in said nucleic acid molecule using a minimally disruptive conversion method, thereby producing a converted nucleic acid; c) amplifying the converted nucleic acid by polymerase chain reaction, thereby producing an amplified converted nucleic acid; d) probing the amplified converted nucleic acid with nucleic acid probes complementary to a pre-identified panel of CpG or CH loci to enrich for sequences corresponding to the pre-identified panel of CpG or CH loci, thereby producing probed converted nucleic acid; e) determining the nucleic acid sequence of the probed converted nucleic acid to a depth of greater than 100x; f) comparing the nucleic acid sequences of said probed converted nucleic acids with reference nucleic acid sequences of said pre-identified panel of CpG or CH loci to obtain a set of measured input features representative of the methylation profiles of healthy and cancer subjects; g) training a machine learning model to generate a classifier that distinguishes between healthy and cancer subjects.
27. 27. The method of claim 26, wherein the pre-identified panel of CpG or CH loci comprises loci associated with transcription start sites.
28. 27. The method of claim 26, further comprising determining hemimethylated CpG or CH loci.
29. 27. The method of claim 26, further comprising identifying the tissue origin of the nucleic acid molecule.
30. 27. The method of claim 26, further comprising identifying the genomic location and fragment length of the nucleic acid molecule.
31. 27. The method of claim 26, wherein the input features are selected from: base-by-base methylation % for CpG, base-by-base methylation % for CHG, base-by-base methylation % for CHH, number or percentage of observed fragments differing in the number or percentage of methylated CpGs in a region, conversion efficiency, hypomethylated blocks, CPG methylation level, CHH methylation level, CHG methylation level, fragment length, fragment midpoint, chrM methylation level, LINE1 methylation level, ALU methylation level, dinucleotide coverage, coverage uniformity, overall average CpG coverage, and average coverage at CpG islands, CGI shelves, and CGI shores.
32. 1. A classifier for distinguishing between healthy and cancer populations of individuals, the classifier comprising a set of measurements representing methylation profiles from methylation sequencing data of each of healthy and cancer subjects, the measurements being used to generate a set of features corresponding to characteristics of the methylation profiles, the set of features being input into a machine learning or statistical model, the machine learning or statistical model providing a feature vector useful as a classifier for distinguishing between healthy and cancer populations of individuals.
33. 1. A method for detecting cancer in a population of subjects, the method comprising: a) assaying nucleic acids of a biological sample from a subject using minimally destructive targeted conversion methyl sequencing to obtain a methylation profile of the nucleic acid; b) classifying said biological samples by inputting said methylation profile into a training algorithm that classifies each sample between healthy and cancerous subjects; and c) if the training algorithm classifies the biological sample as negative for cancer with a particular confidence value, outputting a report on a computer screen identifying the biological sample as negative for cancer.
34. 34. The method of claim 33, wherein the cancer is colon cancer.
35. 1. A system for classifying individuals based on methylation status, the system comprising: a) a computer readable medium product comprising a classifier, the classifier comprising a set of measurements representing methylation profiles from methylation sequencing data of each of healthy and cancer subjects, the measurements being used to generate a set of features corresponding to characteristics of each of the methylation profiles of the healthy and cancer subjects, the set of features being input into a machine learning or statistical model, the machine learning or statistical model providing a feature vector useful as a classifier for distinguishing between populations of healthy and cancer individuals; b) one or more processors for executing instructions stored on said computer-readable media product.
36. 36. The system of claim 35, wherein the system comprises a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first order polynomial kernel support vector machine classifier, a second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal problem optimization algorithm classifier, a naive Bayes algorithm classifier, and a non-negative matrix factorization (NMF) prediction algorithm classifier.
37. 36. The system of claim 35, wherein the system comprises means for performing any of the above methods.
38. 36. The system of claim 35, wherein the system comprises one or more processors configured to perform any of the above methods.
39. 36. The system of claim 35, wherein the system comprises modules for performing each of the steps of the method.
40. 1. A method for monitoring minimal residual disease status in a subject who has previously been treated for a disease, the method comprising: a) determining a baseline methylation profile of a biological sample obtained from a subject at baseline methylation status; b) determining a test methylation profile of a biological sample obtained from the subject at one or more predetermined time points after the baseline methylation status; and c) determining a change in the test methylation profile as compared to a baseline methylation profile, said change indicating a change in the minimal residual disease status of the subject.
41. 41. The method of claim 40, wherein the minimal residual disease status is selected from response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.
42. 41. The method of claim 40, wherein the disease is colon cancer.
43. A tumor detection kit comprising reagents for carrying out the above method and instructions for detecting tumor signals.
44. 44. The kit for tumor detection of claim 43, wherein the reagents are selected from the group consisting of primer sets, PCR reaction components, sequencing reagents, minimally destructive conversion reagents, and library preparation reagents.