Compositions and methods for improved resolution of 5-hydroxymethylated cytosine in nucleic acid sequencing

JP2024528704A5Pending Publication Date: 2025-07-08FREENOM HLDG INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024503898
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-20
Filing Date
2022-07-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing sequencing methods for detecting 5-hydroxymethylated cytosines (5hmC) suffer from limitations such as lack of nucleotide resolution, false positives, high sample input requirements, and indirect readouts, which affect the accuracy of disease diagnosis and classification models.

Method used

The use of modified adapters containing 5hmC, 5-(β-glucosyloxymethyl)cytosine (5gmC), and 5-carboxycytosine (5caC) nucleotides during nucleic acid sequencing, which are ligated to nucleic acid fragments to improve hydroxymethylation sequence information, and conversion of unmethylated and methylated cytosines to uracils without converting 5hmC, followed by sequencing to obtain accurate hydroxymethylation status data.

Benefits of technology

Enhances the resolution and accuracy of hydroxymethylation detection, enabling more precise disease diagnosis and classification models with improved sensitivity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides oligonucleotide adaptor compositions, methods and systems for improving the resolution of 5hmC sequencing, which are useful for improving the quality of nucleic acid sequencing libraries and nucleic acid methylation profiling. Methods for applying the improved oligonucleotide adaptors and sequencing methods for machine learning classifier generation, as well as methods for detecting cell proliferation disorders such as cancer, are also provided. Methods for applying the improved oligonucleotide adaptors, as well as methods for applying target nucleic acid enrichment, and sequencing methods for improving the quality of nucleic acid sequencing libraries and nucleic acid methylation profiling are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 223,661, filed July 20, 2021, which is incorporated by reference in its entirety.

[0002] The present disclosure relates generally to improved adaptors and methods for performing methylation analysis of nucleic acid sequences. The present disclosure relates to sequencing adaptors and methods of use for improving sequencing resolution of 5-hydroxymethylated cytosines, which may be useful for nucleic acid methylation pattern analysis. [Background technology]

[0003] DNA methylation occurs primarily at cytosines in CpG dinucleotides and acts as an epigenetic mark with a functional role in gene regulation. Methylation marks are heritable and their genome-wide profiles vary from tissue to tissue. In cancer, gene-specific methylation profiles can be aberrant but retain similarity to the tissue of origin, making methylation marks useful biomarkers for cancer diagnosis and prognosis.

[0004] 5-Methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) are two forms of epigenetic modifications at the 5-carbon position of cytosine that are associated with gene silencing and gene activation, respectively. These methylation marks provide various types of information that can be used to build classification models for inferring the presence of cancer. Quality sequence information is desirable to generate classification models for inferring disease with high sensitivity and specificity, and such information can be lost during sample processing and sequencing, thereby affecting the accuracy of such models.

[0005] Although several sequencing methods can be used to identify 5hmC, these methods have advantages and disadvantages that impact their adoption for commercial screening and diagnostic applications, such as lack of nucleotide resolution, false positive 5hmC calls, high sample input requirements, inference by subtraction rather than direct readout, and the quality of the sequencing libraries that are generated for sequencing from nucleic acid samples. Thus, tools and methods may be needed to improve the quality of hydroxymethylation status information provided from nucleic acid sequencing that may be useful in disease diagnosis, prognosis, and classification models of progression. Summary of the Invention

[0006] The present disclosure provides compositions, methods, and systems directed to improved detection of hydroxymethylated cytosine during nucleic acid sequencing.The methods and compositions used in such methods described herein can be used to overcome the limitations of unmethylated and methylated cytosine conversion methods, such as TAB-seq and ACE-seq, used prior to nucleic acid sequencing.In various aspects, using modified adaptors containing 5hmC, or a combination of 5-(β-glucosyloxymethyl)cytosine (5gmC) and 5-carboxycytosine (5caC) or 5-carboxymethylcytosine (5cxmC), and ligating such adaptors to nucleic acid fragments in biological samples can improve the resolution of hydroxymethylated sequence information in samples.

[0007] In an aspect, the present disclosure provides an oligonucleotide adaptor, which comprises one or more 5hmC, 5gmC, 5caC, 5cxmC nucleotides, or combinations thereof, but does not comprise a cytosine nucleotide, and can be used in ligation to a nucleic acid molecule in a biological sample for nucleic acid sequencing. In some embodiments, the cytosine nucleotide is not present in the flow cell binding region or primer binding site of the adaptor. In some embodiments, the cytosine nucleotide is present in the UMI portion of the adaptor, but is not present in the non-UMI portion of the adaptor. In some embodiments, the cytosine nucleotide is present in the primer binding site portion of the adaptor, but is not present in the non-primer binding site portion of the adaptor. The oligonucleotide can be ligated to a nucleic acid sequence before being treated with the necessary conditions to convert unmethylated and methylated cytosines in the nucleic acid sequence to uracil, and can be hybridized to a primer for downstream amplification and sequencing methods.

[0008] In another aspect, the disclosure provides a method for providing hydroxymethylation status data of a nucleic acid in a biological sample, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) applying to at least a portion of the ligated nucleic acid or derivative thereof conversion conditions that convert unmethylated and methylated cytosine nucleotides of the ligated nucleic acid to uracil nucleotides, but do not convert hydroxymethylated cytosine nucleotides to uracil nucleotides, thereby producing a converted nucleic acid; and d) sequencing at least a portion of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid and provide hydroxymethylation status data of the nucleic acid.

[0009] In some embodiments, the oligonucleotide adaptor does not include a cytosine nucleotide in the flow cell binding region or the primer binding site of the oligonucleotide adaptor.

[0010] In some embodiments, the method further comprises, after b) or before c), subjecting at least a portion of the ligated nucleic acid to glucosylation with β-glucosyltransferase (β-GT) / UDP-glucose to convert 5hmC nucleotides to 5gmC nucleotides.

[0011] In some embodiments, the converting conditions include bisulfite treatment, enzymatic treatment, or a combination thereof.

[0012] In some embodiments, the oligonucleotide adaptor comprises a 5hmC nucleotide.

[0013] In some embodiments, the oligonucleotide adaptor comprises 5gmC and 5caC nucleotides.

[0014] In some embodiments, the oligonucleotide adaptors comprise 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or combinations thereof.

[0015] In some embodiments, the converting conditions include treatment with β-GT, a cytosine dioxygenase enzyme, a carboxymethyltransferase, an apolipoprotein B mRNA editing catalytic polypeptide-like protein (AID / APOBEC), or a combination thereof.

[0016] In some embodiments, the cytosine dioxygenase enzyme comprises ten eleven translocation protein 1 (TET1), ten eleven translocation protein 2 (TET2), ten eleven translocation protein 3 (TET3), or a functional variant thereof.

[0017] In some embodiments, the method further comprises treating the oligonucleotide adaptor with a TET enzyme after a) or before b).

[0018] In some embodiments, the method further comprises performing sequence enrichment after b) or before c).

[0019] In some embodiments, sequence enrichment comprises target capture hybridization.

[0020] In some embodiments, at least a portion of the ligated nucleic acids is amplified prior to sequencing.

[0021] In some embodiments, the method further comprises amplifying at least a portion of the ligated nucleic acids prior to sequencing.

[0022] In some embodiments, the method further comprises preparing a nucleic acid sequencing library prior to the amplification.

[0023] In some embodiments, the method further comprises aligning the nucleic acid sequence to a reference genome.

[0024] In some embodiments, oligonucleotide adaptors are chemically synthesized using 5hmC phosphoramidites.

[0025] In some embodiments, the oligonucleotide adaptor comprises a 5gmC nucleotide and a 5caC nucleotide, and the oligonucleotide adaptor is generated at least in part by synthesizing a 5mC-containing oligonucleotide using phosphoramidite chemistry and enzymatically treating the 5mC-containing oligonucleotide with a TET enzyme and β-GT / UDP-glucose.

[0026] In some embodiments, oligonucleotide adaptors are synthesized using terminal deoxynucleotidyl transferase (TdT)-mediated enzymatic oligonucleotide synthesis.

[0027] In some embodiments, the method further comprises methylating unmethylated cytosine nucleotides in the 5mC-containing oligonucleotide using a SAM-dependent C5-methyltransferase (C5-MT) or another DNA cytosine-5 methyltransferase.

[0028] In some embodiments, the method further comprises ligating an oligonucleotide adaptor to at least a portion of the nucleic acid isolated from the biological sample.

[0029] In some embodiments, oligonucleotide adaptors are synthesized using enzymatic oligonucleotide synthesis techniques.

[0030] In some embodiments, the biological sample comprises cell-free DNA (cfDNA).

[0031] In some embodiments, the nucleic acid is cfDNA.

[0032] In some embodiments, the biological sample is obtained or derived from an individual, and the hydroxymethylation status data is associated with an abnormal cellular condition or disease and provides a classification of the individual having the abnormal cellular condition or disease.

[0033] In some embodiments, the abnormal cellular condition or disease is stage 1 cancer, stage 2 cancer, stage 3 cancer, or stage 4 cancer.

[0034] In some embodiments, the oligonucleotide adaptor comprises a unique molecular identifier.

[0035] In some embodiments, the biological sample is selected from the group consisting of bodily fluids, feces, colonic effluent, urine, cerebrospinal fluid, plasma, serum, whole blood, isolated blood cells, cells isolated from the blood, and combinations thereof.

[0036] In some embodiments, the method optionally includes characterizing the hydroxymethylation status data and processing the characterized hydroxymethylation status data using a machine learning model trained to classify biological samples into groups according to pre-specified or pre-selected biological characteristics.

[0037] In some embodiments, the characterized hydroxymethylation status data corresponds to a characteristic of a nucleic acid sequence in a biological sample.

[0038] In some embodiments, the nucleic acid sequence characteristic is selected from the presence or absence of precancer, cancer or stage of cancer, or prognosis of cancer in the subject.

[0039] In another aspect, the present disclosure provides a method for generating an oligonucleotide adaptor, the method comprising: a) synthesizing, at least in part, a 5mC-containing oligonucleotide by phosphoramidite chemistry; b) contacting the 5mC-containing oligonucleotide with a TET enzyme and β-GT / UDP-glucose to convert the 5mC nucleotide to a 5gmC nucleotide or a 5caC nucleotide, thereby generating an oligonucleotide adaptor.

[0040] In some embodiments, oligonucleotide adaptors are synthesized using terminal deoxynucleotidyl transferase (TdT)-mediated enzymatic oligonucleotide synthesis.

[0041] In some embodiments, the oligonucleotide adaptor comprises 5gmC and 5caC nucleotides.

[0042] In some embodiments, the method further comprises methylating unmethylated cytosine nucleotides in the 5mC-containing oligonucleotide using a SAM-dependent C5-methyltransferase (C5-MT), or another DNA cytosine-5 methyltransferase.

[0043] In some embodiments, the method further comprises ligating an oligonucleotide adaptor to at least a portion of the nucleic acid isolated from the biological sample.

[0044] In another aspect, the disclosure provides a method for generating an oligonucleotide adaptor, the method comprising synthesizing, at least in part, an oligonucleotide containing 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or combinations thereof by phosphoramidite chemistry, thereby generating the oligonucleotide adaptor.

[0045] In some embodiments, oligonucleotide adaptors are synthesized using enzymatic oligonucleotide synthesis techniques.

[0046] In some embodiments, the method further comprises ligating an oligonucleotide adaptor to at least a portion of the nucleic acid isolated from the biological sample.

[0047] In another aspect, the present disclosure provides a method for training a machine learning model to generate a hydroxymethylation profile of a nucleic acid in a biological sample, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) applying conversion conditions to at least a portion of the ligated nucleic acids that convert unmethylated and methylated cytosine nucleotides in the ligated nucleic acids to uracil nucleotides, thereby generating converted nucleic acids; d) sequencing at least a portion of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid and provide hydroxymethylation status data of the nucleic acid; e) using the hydroxymethylation status data to train a machine learning model to generate a hydroxymethylation profile; Includes.

[0048] In some embodiments, e) further comprises characterizing the hydroxymethylation status data. In some embodiments, the oligonucleotide adaptor does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the oligonucleotide adaptor.

[0049] In some embodiments, the method further comprises, after b) or before c), at least a portion of the ligated nucleic acid is glucosylated, at least in part, with β-GT / UDP-glucose to convert 5hmC nucleotides to 5gmC nucleotides.

[0050] In some embodiments, the biological sample comprises cell-free DNA (cfDNA).

[0051] In another aspect, the present disclosure provides a method for determining a hydroxymethylation profile of cfDNA in a biological sample obtained or derived from an individual, the method comprising: a) obtaining a biological sample containing cfDNA; b) ligating an oligonucleotide adaptor to at least a portion of the cfDNA in the biological sample, thereby generating a ligated cfDNA, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) applying conversion conditions to at least a portion of the ligated cfDNA or derivative thereof that convert unmethylated and methylated cytosine nucleotides in the ligated cfDNA to uracil nucleotides, thereby producing converted cfDNA; d) sequencing at least a portion of the converted cfDNA to obtain a nucleic acid sequence of the converted cfDNA and provide hydroxymethylation status data of the cfDNA; e) aligning the nucleic acid sequence of the converted cfDNA to a reference nucleic acid sequence to determine the hydroxymethylation profile of the biological sample.

[0052] In some embodiments, the method further comprises amplifying the ligated cfDNA prior to sequencing.

[0053] In some embodiments, the method further comprises preparing a nucleic acid sequencing library prior to the amplification.

[0054] In some embodiments, the oligonucleotide adaptor does not include a cytosine nucleotide in the flow cell binding region or in the primer binding site in the oligonucleotide adaptor.

[0055] In some embodiments, the method further comprises, after b) or before c), glucosylating at least a portion of the ligated cfDNA, at least in part, with β-GT / UDP-glucose to convert hydroxymethylated cytosine nucleotides to 5gmC nucleotides.

[0056] In some embodiments, the hydroxymethylation profile is associated with an abnormal cellular condition or disease and provides a classification of individuals having the abnormal cellular condition or disease.

[0057] In some embodiments, the abnormal cellular condition or disease is stage 1 cancer, stage 2 cancer, stage 3 cancer, or stage 4 cancer.

[0058] In some embodiments, the oligonucleotide adaptor comprises a unique molecular identifier.

[0059] In some embodiments, the converting conditions include using chemical methods, enzymatic methods, or a combination thereof.

[0060] In some embodiments, the converting conditions include treatment with bisulfite, hydrogen sulfite, disulfite, or a combination thereof.

[0061] In some embodiments, the biological sample is selected from the group consisting of bodily fluids, feces, colonic effluent, urine, cerebrospinal fluid, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.

[0062] In another aspect, the present disclosure provides a method for generating a classifier for a biological sample, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) applying conversion conditions to at least a portion of the ligated nucleic acid that convert unmethylated and methylated cytosine nucleotides in the ligated nucleic acid to uracil nucleotides, thereby producing a converted nucleic acid; d) sequencing at least a portion of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid and provide hydroxymethylation status data of the nucleic acid; e) using the hydroxymethylation status data to train a machine learning model to generate a classifier; Includes.

[0063] In some embodiments, the oligonucleotide adaptor does not include a cytosine nucleotide in the flow cell binding region or in the primer binding site in the oligonucleotide adaptor.

[0064] In some embodiments, the method includes after b) or before c) glucosylating at least a portion of the ligated nucleic acid, at least in part, with β-GT / UDP-glucose to convert hydroxymethylated cytosine nucleotides to 5gmC nucleotides.

[0065] In another aspect, the disclosure provides a method for generating a classifier for a biological sample obtained or derived from an individual, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof, and no cytosine nucleotides; c) applying conversion conditions to at least a portion of the ligated nucleic acid that convert unmethylated and methylated cytosine nucleotides in the ligated nucleic acid to uracil nucleotides, thereby generating a converted nucleic acid; d) sequencing at least a portion of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid and provide hydroxymethylation status data of the nucleic acid; e) using the hydroxymethylation status data to train a machine learning model to generate a classifier; Includes.

[0066] In another aspect, the disclosure provides a method for detecting a cell proliferative disorder in a subject, the method comprising: a) obtaining a biological sample comprising nucleic acid from a subject; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) providing conversion conditions to at least a portion of the ligated nucleic acids that convert unmethylated uracil and methylated cytosine nucleotides in the ligated nucleic acids to uracil nucleotides, thereby producing converted nucleic acids; d) sequencing at least a portion of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid and provide hydroxymethylation status data of the nucleic acid; e) processing the hydroxymethylation status data using a machine learning model trained to distinguish between healthy subjects and subjects with a cell proliferative disorder to provide an output value related to the presence or susceptibility to a cell proliferative disorder, thereby indicating the presence or susceptibility of a cell proliferative disorder in the subject; Includes.

[0067] In some embodiments, the adaptor does not include a cytosine nucleotide in the flow cell binding region or in the primer binding site in the oligonucleotide adaptor.

[0068] In some embodiments, the method further comprises, after b) or before c), subjecting at least a portion of the ligated nucleic acid to glucosylation, at least in part, with β-GT / UDP-glucose, to convert hydroxymethylated cytosine nucleotides to 5gmC nucleotides.

[0069] In some embodiments, the cell proliferative disorder comprises colorectal cancer, breast cancer, ovarian cancer, prostate cancer, lung cancer, pancreatic cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, thyroid cancer, or bladder cancer.

[0070] In some embodiments, the machine learning model is tuned to detect a cell proliferative disorder with a preselected sensitivity and specificity.

[0071] In some embodiments, the machine learning model classifies the presence or susceptibility to a cell proliferative disorder with at least about 80% sensitivity.

[0072] In some embodiments, the converting conditions include bisulfite treatment, enzymatic treatment, or a combination thereof.

[0073] In some embodiments, the oligonucleotide adaptor contains a 5hmC nucleotide in place of a cytosine nucleotide in the flow cell binding region or primer binding site in the oligonucleotide adaptor.

[0074] In some embodiments, the oligonucleotide adaptors comprise a mixture of 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or combinations thereof.

[0075] In some embodiments, the converting conditions include treatment with β-GT, a cytosine dioxygenase enzyme, a carboxymethyltransferase, AID / APOBEC, or a combination thereof.

[0076] In some embodiments, the cytosine dioxygenase enzyme comprises TET1, TET2, TET3, or a functional variant thereof.

[0077] In some embodiments, the method further comprises treating the oligonucleotide adaptor with a TET enzyme after a) or before b).

[0078] In some embodiments, the method further comprises performing sequence enrichment after b) or before c).

[0079] In some embodiments, sequence enrichment comprises target capture hybridization.

[0080] In some embodiments, the method further comprises amplifying at least a portion of the ligated nucleic acids prior to sequencing.

[0081] In some embodiments, the method further comprises aligning the nucleic acid sequence to a reference genome.

[0082] In some embodiments, the method further comprises characterizing the hydroxymethylation status data and processing the characterized hydroxymethylation status data using a machine learning model trained to classify biological samples into groups by pre-specified or pre-selected biological characteristics.

[0083] In some embodiments, the characterized hydroxymethylation status data corresponds to a characteristic of a nucleic acid sequence in a biological sample.

[0084] In some embodiments, the nucleic acid sequence characteristic is selected from the presence or absence of precancer, cancer or stage of cancer, or prognosis of cancer in the subject.

[0085] In another aspect, the disclosure provides a method for monitoring minimal residual disease in a subject previously treated for a disease, the method comprising determining a hydroxymethylation profile as a baseline hydroxymethylation state and further determining the hydroxymethylation profile at each of one or more predetermined time points, wherein a change in the hydroxymethylation profile from the baseline hydroxymethylation state indicates a change in minimal residual disease status at the baseline hydroxymethylation state in the subject.

[0086] In some embodiments, minimal residual disease is indicated by response to treatment, tumor load, residual tumor after surgery, recurrence, secondary screen, primary screen, or cancer progression.

[0087] In some embodiments, the method further comprises determining the subject's response to the treatment.

[0088] In some embodiments, the method further comprises monitoring tumor burden in the subject.

[0089] In some embodiments, the method further comprises detecting residual tumor in the subject after surgery.

[0090] In some embodiments, the method further comprises detecting recurrence in the subject.

[0091] In some embodiments, the method is performed as a secondary screen for the subject.

[0092] In some embodiments, the method is performed as a primary screen for subjects.

[0093] In some embodiments, the method further comprises monitoring the progression of the cancer in the subject.

[0094] In another aspect, the disclosure provides a non-transitory computer readable medium comprising stored instructions that, when executed by one or more processors, is operable to implement a classifier for classifying a subject as having or not having a cell proliferative disorder based on hydroxymethylation status data obtained from a nucleic acid library generated using oligonucleotide adaptors ligated to nucleic acids in a biological sample, wherein the oligonucleotide adaptors comprise 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or combinations thereof.

[0095] In some embodiments, the oligonucleotide adaptor does not include a cytosine nucleotide in the flow cell binding region or in the primer binding site in the oligonucleotide adaptor.

[0096] In some embodiments, the classifier for detecting a cell proliferative disorder is further configured to determine a tissue of origin of the cell proliferative disorder.

[0097] In some embodiments, the classifier is trained using a training vector obtained from training biological samples, a first subset of the training biological samples being identified as having a cell proliferative disorder and a second subset of the training biological samples being identified as not having a cell proliferative disorder.

[0098] In another aspect, the disclosure provides a method of sequencing a nucleic acid to provide hydroxymethylation status data of a nucleic acid molecule in a biological sample, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating an oligonucleotide adaptor to at least a portion of a nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the adaptor comprises 5hmC nucleotides, 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or a combination thereof; c) applying to at least a portion of the ligated nucleic acid conversion conditions necessary to convert unmethylated and methylated cytosines, but not hydroxymethylated cytosines, in the nucleic acid to uracil; and d) sequencing the nucleic acid to obtain a nucleic acid sequence of the nucleic acid and provide hydroxymethylation status data for the nucleic acid molecule.

[0099] In some embodiments, the adaptor does not contain a cytosine nucleotide in the flow cell binding region or the primer binding site of the adaptor.

[0100] In some embodiments, the method includes, after the ligation procedure, glucosylating the ligated nucleic acid with β-GT / UDP-glucose to convert 5hmC nucleotides to 5gmC nucleotides.

[0101] In some embodiments, the conversion conditions include bisulfite treatment, enzyme treatment, or a combination of both.

[0102] In some embodiments, the oligonucleotide adaptor comprises all 5hmC nucleotides in place of cytosine nucleotides in the designed oligonucleotide adaptor sequence.

[0103] In some embodiments, the oligonucleotide adaptors comprise a mixture of 5gmC, 5caC, and / or 5cxmC nucleotides in place of cytosine nucleotides in the designed oligonucleotide adaptor sequence.

[0104] In some embodiments, the enzymatic treatment includes treatment with one or more of β-glucosyltransferase (β-GT), a cytosine dioxygenase enzyme (such as TET1, TET2, TET3, or a functional variant thereof), carboxymethyltransferase, or AID / APOBEC.

[0105] In some embodiments, a sequence enrichment operation is performed after operation b) or before operation c).

[0106] In some embodiments, the sequence enrichment procedure is target capture hybridization.

[0107] In some embodiments, the ligated nucleic acid is amplified prior to sequencing.

[0108] In some embodiments, the nucleic acid sequences obtained from the sequencing are aligned to a reference genome.

[0109] In some embodiments, 5hmC-containing adapter oligonucleotides can be chemically synthesized using 5-hydroxymethyl-modified cytidine phosphoramidites.

[0110] In some embodiments, adapter oligonucleotides containing a mixture of 5gmC and 5caC can be generated by first synthesizing 5mC-containing adapters using phosphoramidite chemistry and then enzymatically treating them with TET enzyme + β-GT / UDP-glucose.

[0111] 1. A method for manufacturing oligonucleotide sequencing adaptors, the method comprising: a) synthesizing an oligonucleotide containing 5mC by phosphoramidite chemistry; b) converting the oligonucleotide with TET enzyme + β-GT / UDP-glucose under conditions sufficient to oxidize the oligonucleotide at the 5mC nucleotide; c) ligating the oxidized oligonucleotide to a polynucleic acid molecule isolated from the biological sample; Includes.

[0112] In some embodiments, 5hmC-containing adaptors can be directly synthesized using enzymatic oligonucleotide synthesis using terminal deoxynucleotidyl transferase (TdT)-mediated enzymatic oligo synthesis.

[0113] In some embodiments, adaptors containing a mixture of 5gmC and 5caC can be generated by first synthesizing 5mC-containing adaptors using enzymatic oligonucleotide synthesis techniques and then enzymatically treating them with TET enzyme plus β-GT / UDP-glucose.

[0114] In some embodiments, adapters containing 5mC can be generated by methylating adapters containing unmethylated cytosines using SAM-dependent C5-methyltransferase (C5-MT), or other DNA cytosine-5 methyltransferases.

[0115] 1. A method for producing an oligonucleotide sequencing adaptor, the method comprising: a) synthesizing oligonucleotides containing 5gmC, 5caC, and / or 5cxmC by phosphoramidite chemistry; b) ligating the synthesized oligonucleotide to a polynucleic acid molecule isolated from a biological sample; Includes.

[0116] In some embodiments, the 5caC-containing adaptor can be directly synthesized using enzymatic oligonucleotide synthesis techniques.

[0117] In another aspect, a method is provided for generating a hydroxymethylation profile of a biological sample obtained or derived from an individual, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating oligonucleotide adaptors to nucleic acids in the biological sample, wherein the adaptors comprise 5hmC, 5gmC, 5caC, 5cxmC, or combinations thereof, and no cytosine nucleotides; c) subjecting the ligated nucleic acid to conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil; d) sequencing the nucleic acid to obtain a nucleic acid sequence of the nucleic acid and provide hydroxymethylation status data for the nucleic acid; e) characterizing the hydroxymethylation status data and training a machine learning model to generate a methylation profile using the hydroxymethylation status data; Includes.

[0118] In some embodiments, the adaptor comprises 5hmC, 5gmC, 5caC, 5cxmC, or a combination thereof, and does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the adaptor.

[0119] In some embodiments, the method comprises glucosylating the ligated nucleic acid with β-GT / UDP-glucose to convert 5hmC to 5gmC, followed by applying conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil.

[0120] In some embodiments, the nucleic acid sample is a cell-free DNA (cfDNA) sample.

[0121] In another aspect, the present disclosure provides a method for determining a hydroxymethylation profile of a cfDNA sample obtained or derived from an individual, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating oligonucleotide adaptors to nucleic acids in the biological sample, wherein the adaptors comprise 5hmC, 5gmC, 5caC, 5cxmC, or combinations thereof, and no cytosine nucleotides; c) subjecting the ligated nucleic acid to conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid of the biological sample to uracil; d) sequencing the nucleic acid to obtain a nucleic acid sequence of the nucleic acid and provide hydroxymethylation status data for the nucleic acid; e) aligning the nucleic acid sequence of the converted nucleic acid molecule to a reference nucleic acid sequence to determine the hydroxymethylation profile of the individual. Includes.

[0122] In some embodiments, a nucleic acid sequencing library is prepared prior to amplification.

[0123] In some embodiments, the adaptor comprises 5hmC, 5gmC, 5caC, 5cxmC, or a combination thereof, and does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the adaptor.

[0124] In some embodiments, the reference nucleic acid sequence is a reference genome.

[0125] In some embodiments, the method includes glucosylating the ligated nucleic acid with β-GT / UDP-glucose to convert 5hmC to 5gmC, followed by applying conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil.

[0126] In some embodiments, the hydroxymethylation profile is associated with an abnormal cellular condition or disease and provides for classification of the subject as having an abnormal cellular condition or disease.

[0127] In some embodiments, an oligonucleotide adaptor comprising a unique molecular identifier is ligated to unconverted nucleic acid in the cfDNA sample prior to a).

[0128] In some embodiments, the nucleic acid molecule is subjected to cytosine to uracil conversion conditions using chemical methods, enzymatic methods, or a combination thereof.

[0129] In some embodiments, the cfDNA in the biological sample is treated with bisulfite, bisulfite, disulfite, or a combination thereof.

[0130] In some embodiments, the biological sample obtained from the subject contains the nucleic acid molecule and is a bodily fluid, feces, colonic effluent, urine, cerebrospinal fluid, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, or a combination thereof.

[0131] In some embodiments, the cell proliferative disorder is selected from stage 1 cancer, stage 2 cancer, stage 3 cancer, and stage 4 cancer.

[0132] In another aspect, a method is provided for generating a classifier for a nucleic acid sample obtained or derived from an individual, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating oligonucleotide adaptors to nucleic acids in the biological sample, wherein the adaptors comprise 5hmC, 5gmC, 5caC, 5cxmC, or combinations thereof, and no cytosine nucleotides; c) subjecting the ligated nucleic acid to conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil; d) sequencing the nucleic acid to obtain a nucleic acid sequence of the nucleic acid and provide hydroxymethylation status data for the nucleic acid; e) using the hydroxymethylation status data to train a machine learning model to generate a classifier; Includes.

[0133] In some embodiments, the adaptor comprises 5hmC, 5gmC, 5caC, 5cxmC, or a combination thereof, and does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the adaptor.

[0134] In some embodiments, the method comprises glucosylating the ligated nucleic acid with β-GT / UDP-glucose to convert hydroxymethylated C to 5gmC, followed by applying conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil.

[0135] In another aspect, the disclosure provides a method for detecting a cell proliferative disorder in a subject, the method comprising: a) obtaining a biological sample containing nucleic acid; b) ligating oligonucleotide adaptors to nucleic acids in the biological sample, wherein the adaptors comprise 5hmC, 5gmC, 5caC, 5cxmC, or combinations thereof, and no cytosine nucleotides; c) subjecting the ligated nucleic acid to conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil; d) sequencing the nucleic acid to obtain a nucleic acid sequence of the nucleic acid and provide hydroxymethylation status data for the nucleic acid; f) processing the hydroxymethylation status data using a machine learning model trained to distinguish between healthy subjects and subjects with a cell proliferative disorder and providing an output value related to the presence of a cell proliferative disorder, thereby indicating the presence of a cell proliferative disorder in the subject; Includes.

[0136] In some embodiments, the adaptor comprises 5hmC, 5gmC, 5caC, 5cxmC, or a combination thereof, and does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the adaptor.

[0137] In some embodiments, the method comprises glucosylating the ligated nucleic acid with β-GT / UDP-glucose to convert hydroxymethylated C to 5gmC, followed by applying conversion conditions necessary to convert unmethylated and methylated cytosines in the nucleic acid to uracil.

[0138] In various embodiments, the different types of cell proliferative disorders are selected from colorectal cancer, breast cancer, ovarian cancer, prostate cancer, lung cancer, pancreatic cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, thyroid cancer, or bladder cancer.

[0139] In some embodiments, the machine learning classifier is tuned to provide preselected sensitivity and specificity for different types of cell proliferative disorders to be detected according to the need for cancer diagnosis and confirmatory diagnosis of the cell proliferative disorder being colorectal cancer, breast cancer, ovarian cancer, prostate cancer, lung cancer, pancreatic cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, thyroid cancer, or bladder cancer, or combinations thereof.

[0140] In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with at least about 80% sensitivity. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with at least about 90% sensitivity. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with at least about 95% sensitivity. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a positive predictive value (PPV) of at least about 70%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a PPV of at least about 80%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a PPV of at least about 90%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a PPV of at least about 95%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a PPV of at least about 99%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with a negative predictive value (NPV) of at least about 80%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with an NPV of at least about 90%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer with an NPV of at least about 95%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer in a subject with an NPV of at least about 99%. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer in a subject with an area under the curve (AUC) of at least about 0.90. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer in a subject with an AUC of at least about 0.95. In some embodiments, the machine learning model classifies the presence or susceptibility of cancer in a subject with an AUC of at least about 0.99.

[0141] In some embodiments, the conversion conditions include bisulfite treatment, enzyme treatment, or a combination of both.

[0142] In some embodiments, the oligonucleotide adaptor comprises all 5hmC nucleotides in place of cytosine nucleotides in the flow cell binding region, and optionally also comprises a primer binding site in the adaptor in a given oligonucleotide adaptor sequence.

[0143] In some embodiments, the oligonucleotide adaptors comprise a mixture of 5gmC and 5caC, or 5cxmC and cytosine nucleotides in the designed oligonucleotide adaptor sequence.

[0144] In some embodiments, the enzymatic treatment includes treatment with one or more of β-glucosyltransferase (β-GT), a cytosine dioxygenase enzyme (such as TET1, TET2, TET3, or a functional variant thereof), carboxymethyltransferase, or AID / APOBEC.

[0145] In some embodiments, the use of TET enzyme enzymatic treatment is performed on the adaptors prior to ligation.

[0146] In some embodiments, a sequence enrichment operation is performed after operation b) or before operation c).

[0147] In some embodiments, the sequence enrichment procedure is target capture hybridization.

[0148] In some embodiments, the ligated nucleic acid is amplified prior to sequencing.

[0149] In some embodiments, the nucleic acid sequences obtained from the sequencing are aligned to a reference genome.

[0150] In some embodiments, the hydroxymethylation status data is characterized and processed using a trained machine learning model that is trained to classify samples into groups according to pre-specified or pre-selected biological characteristics.

[0151] In some embodiments, the set of features is identified from the processed nucleic acid sequences using a machine learning model, the set of features can correspond to characteristics of the nucleic acid sequences in the biological sample.

[0152] In some embodiments, the nucleic acid sequence characteristic is selected from the presence or absence of a stage of precancer, cancer, or cancer, or a prognosis of cancer in the individual from which the sample was obtained.

[0153] In another aspect, the disclosure provides a method for monitoring minimal residual disease in a subject previously treated for a disease, the method comprising determining a hydroxymethylation profile as described herein as a baseline hydroxymethylation status and repeating the analysis to determine a hydroxymethylation profile at one or more predefined time points, wherein a change from baseline indicates a change in the minimal residual disease status at baseline in the subject.

[0154] In some embodiments, minimal residual disease is selected from response to treatment, tumor burden, residual tumor after surgery, recurrence, secondary screening, primary screening, and cancer progression.

[0155] In another aspect, a method for determining response to treatment is provided.

[0156] In another aspect, a method for monitoring tumor burden is provided.

[0157] In another aspect, a method for detecting residual tumor after surgery is provided.

[0158] In another aspect, a method for detecting recurrence is provided.

[0159] In another aspect, a method is provided for use as a secondary screen.

[0160] In another aspect, a method is provided for use as a primary screen.

[0161] In another aspect, a method for monitoring the progression of cancer is provided.

[0162] In one aspect, the present disclosure provides a system including a machine learning model classifier for detecting a cell proliferative disorder, the system comprising: a) a computer readable medium comprising a classifier operable to classify a subject as having or not having a cell proliferative disorder based on hydroxymethylation status data obtained from a nucleic acid library generated using oligonucleotide adaptors to nucleic acids in a biological sample, wherein the adaptors comprise 5hmC, 5gmC, 5caC, 5cxmC, or combinations thereof, and no cytosine nucleotides; and b) one or more processors for executing instructions stored on a computer-readable medium; Includes.

[0163] In some embodiments, the adaptor comprises 5hmC, 5gmC, 5caC, 5cxmC, or a combination thereof, and does not comprise a cytosine nucleotide in the flow cell binding region or the primer binding site in the adaptor.

[0164] In some embodiments, the machine learning model classifier for detecting a cell proliferative disorder comprises a tissue of origin determination.

[0165] In some embodiments, the system includes a classifier loaded into a memory of a computer system, a machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having a cell proliferative disorder, and a second subset of the training biological samples identified as not having a cell proliferative disorder.

[0166] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Thus, the drawings and description should be regarded as illustrative in nature, and not as restrictive.

[0167] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned herein are incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over any such conflicting material. [Brief description of the drawings]

[0168] Embodiments of the present disclosure are described, by way of example only, with reference to the accompanying drawings. The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Figure 1A]A schematic diagram showing an exemplary adapter (FIG. 1A) and its use (FIG. 1B) is provided. FIG. 1A provides a generalized example of an adapter used in hydroxymethylation sequencing. The adapter can contain either modified cytosine 5hmC, 5gmC, 5caC, or 5cxmC in the flow cell and primer binding regions. The cytosine in the UMI region may be unmodified or modified with 5mC, 5hmC, 5gmC, 5caC, or 5cxmC. 5m (5-methyl), 5hm (5-hydroxymethyl), 5gm (β-glucosyl-5-hydroxymethyl), 5ca (5-carboxyl), 5cxm (5-carboxymethyl), UMI (unique molecular barcode). [Figure 1B] A schematic diagram showing an exemplary adaptor (FIG. 1A) and its method of use (FIG. 1B) is provided. FIG. 1B provides an example of a process for generating adaptors for hydroxymethylation sequencing. Adaptors can be designed and synthesized using (i) mC nucleotides, or (ii) a combination of 5hmC, 5gmC, 5caC or 5cxmC nucleotides at positions requiring protection from deamination. For process (i), the synthesized adaptors can be oxidized and optionally (*) glucosylated prior to use in ligation. For process (ii), the adaptors are ready for use in ligation. C (cytosine), m (methyl), 5hm (5-hydroxymethyl), 5gm (β-glucosyl-5-hydroxymethyl), 5ca (5-carboxyl), 5cxm (5-carboxymethyl). [Diagram 2] Provides a schematic overview of an exemplary 5hmC-seq assay. The operation of the 5hmC-seq assay starts with an adapter that is protected from downstream enzymatic conversion. The target enrichment operation is optional (*). [Diagram 3] A schematic diagram of a computer system that is programmed or otherwise configured with machine learning models and classifiers to implement the methods provided herein is provided. Detailed Description of the Invention

[0169] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed.

[0170] The present disclosure generally relates to oligonucleotide adaptor compositions useful for sequencing the cytosine hydroxymethylation status of nucleic acids in biological samples. DNA methylation at the 5 carbon position of cytosine (5-methylcytosine; 5mC) is an epigenetic mark that has functional roles in gene silencing, nucleosome positioning, and chromatin organization. In humans, DNA methylation occurs primarily at cytosines in CpG dinucleotides. Methylation marks are heritable and their genome-wide profiles vary from tissue to tissue. In cancer, gene-specific methylation profiles become aberrant but retain similarities to the tissue of origin. These properties make methylation marks highly useful biomarkers for cancer diagnosis and prognosis.

[0171] Circulating cell-free DNA (cfDNA) is released into the blood from dying apoptotic or necrotic cells and thus represents a snapshot of cell death throughout the human body. In tumors, some fraction of cells constantly die and release DNA into the blood circulation as cell-free tumor-derived DNA (ctDNA) fragments. Knowledge of tumor-specific DNA methylation patterns can be utilized as a methylation atlas to interrogate cfDNA and determine whether a given fragment is derived from the tumor or a normal cell type.

[0172] Hydroxymethylation is another epigenetic modification at the 5 carbon position of cytosine (5hmC). This modification may be involved in active demethylation and may play a role in regulating gene expression. In the active demethylation pathway, 5hmC may be generated as the first operation in the iterative oxidation of 5mC. Studies of the genome-wide distribution of 5hmC have demonstrated a dynamic landscape that is strongly associated with gene expression. Changes in 5hmC profiles may be associated with a wide range of disease states, including cell proliferative disorders.

[0173] As used herein, the term "cell proliferative disorder" may generally refer to a disorder or disease involving unregulated or abnormal proliferation of cells. In some non-limiting examples, the disorder is colorectal cell proliferation, prostate cell proliferation, lung cell proliferation, breast cell proliferation, pancreatic cell proliferation, ovarian cell proliferation, uterine cell proliferation, hepatic cell proliferation, esophageal cell proliferation, gastric cell proliferation, or thyroid cell proliferation. In some embodiments, the cell proliferative disorder is colon adenocarcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, prostate cancer, or rectal adenocarcinoma. As used herein, the term "normal" or "healthy" may generally refer to a cell, tissue, plasma, blood, biological sample, or subject that does not have a cell proliferative disorder.

[0174] Improvements in library preparation that capture improved quality hydroxymethylation information of nucleic acids in biological samples may be necessary to increase the sensitivity of classification models and associated clinical screening methods.

[0175] I. Library preparation and adapter ligation for enzymatic hydroxymethylation sequencing Methods are provided for the preparation of sequencing libraries for detecting 5hmC, 5-formylcytosine (5fC), and 5caC in nucleic acid molecules from biological samples. These methods can provide improved library yields and quality, and they are scalable, more manageable, and provide improved adapter protection than other hydroxymethylation sequencing approaches. These methods can also provide base resolution 5hmC data in short-read sequencing, which is more cost-effective and less error-prone than long-read sequencing approaches.

[0176] The method described herein provides a library that is acceptable for DNA hydroxymethylation sequencing applications as well as non-methylation sequencing applications, thereby providing the sequencing data for multiple applications from a single sample.The resulting raw sequencing data can be used for hydroxymethylation status analysis as well as more conventional cfDNA analysis such as copy number alteration, germline variant detection, somatic mutation detection, nucleosome positioning, transcription factor profiling, chromatin immunoprecipitation, etc.

[0177] A. Adapter Ligation for Sequencing Applications In one embodiment, the method can preserve the integrity and information of nucleic acid sequence for hydroxymethylation profiling.In one example, combining 5hmC protection and dsDNA adapter ligation before APOBEC conversion (e.g., deamination) can provide the maximum possible library complexity for library preparation while preserving fragment endpoint information, thereby providing higher sensitivity for detecting rare events such as hydroxymethylated ctDNA.This method can be applied to sample target enrichment or directly applied for genome-wide sequencing.

[0178] Performing adapter ligation prior to 5hmC protection and APOBEC conversion of sample nucleic acid may allow implementation of dsDNA-dependent adapter ligation methods, which maintain end-point information while generating high-complexity libraries. Furthermore, when the fragment length of the sample nucleic acid is small, such as cfDNA (modal size = 167 base pairs, bp), adapter ligation may extend the length of the DNA to approximately twice the length of the adapter (by double-sided ligation), which provides an advantage over unligated cfDNA due to significantly increased recovery efficiency during solid phase reversible immobilization (SPRI)-bead based reaction cleanup operations. Preserving the end-point information of nucleic acid sequences in biological samples may allow more accurate analysis of fragmentation patterns in cfDNA, which may be used as features in machine learning models. To ligate adapter oligonucleotides prior to the protection / conversion workflow process, cytosines in the oligonucleotide adapters that bind to the flow cell surface or the sequencing primer binding site are first modified or protected from deamination during the conversion operation, since C to T substitution during conversion can obstruct sequencing. In some embodiments, this approach can reduce or eliminate the limitations of TAB-seq and ACE-seq by using adapters containing 5hmC, or a mixture of 5gmC and 5caC, at sequence positions where cytosines would normally be located in the adapter design for flow cell attachment and sequencing primer binding. These methods use short-read sequencing, unlike 5hmC-Seal in combination with long-read sequencing, which in some embodiments may be more suitable for the applications discussed herein.

[0179] In some embodiments, 5hmC-containing adaptor oligonucleotides can be directly synthesized using 5-hmC phosphoramidite.After ligation of 5hmC-containing adaptor to cfDNA, the 5hmC nucleotides in the adaptor oligonucleotides and the sample nucleic acid library inserts can be glucosylated using β-glucosyltransferase (β-GT) and substrate, UDP-glucose, during the labeling procedure of hydroxymethylated cytosine.The glucosylation of hydroxymethylated cytosine in the sample nucleic acid can protect the modified cytosine from deamination by subsequent treatment, for example, with bisulfite or APOBEC enzyme.

[0180] In some embodiments, oligonucleotide adaptors containing a mixture of 5gmC and 5caC can be made by first synthesizing 5mC-containing adaptors using phosphoramidite chemistry and then enzymatically treating them with TET enzyme + β-GT / UDP-glucose. Chemical synthesis of 5mC-containing adaptors can be both more efficient with fewer early truncation products and less expensive than chemical synthesis of 5hmC-containing adaptors.

[0181] In some embodiments, 5hmC-containing adaptors can be generated using enzymatic oligonucleotide synthesis techniques. In some embodiments, the enzymatic oligonucleotide synthesis method employs terminal deoxynucleotidyl transferase (TdT), a template-independent polymerase, which attaches supplied deoxynucleotides to the 3'-OH end of DNA.

[0182] In one example, oligonucleotide adaptors can be ligated to the 5' and 3' ends of a population of nucleic acid fragments in a biological sample to generate a sequencing library. In one example, a collection of nucleic acid adaptors is ligated to the nucleic acid fragments in the sample, the collection of adaptors including equal portions of 4bp, 5bp, and 6bp unique molecular identifier (UMI) sequences followed by an invariant thymidine (T) at the last position (e.g., at the 3' end) to allow for T / A overhang ligation. Thus, the UMI can be located adjacent to the library insert nucleic acid. During sequencing, the UMI can also be sequenced as part of the read at the 5' end (alternatively, the UMI can match the library insert at the sequencing read level). The invariant T can be staggered across three positions to maintain base diversity at the sequenced position. In contrast, using a single-length UMI with an invariant thymidine can result in low-complexity sequencing at the position corresponding to the invariant thymidine, resulting in reduced sequencing quality. The first 4bp of each UMI together comprises a set of 4bp core UMI sequences, which have an edit distance of 2 or more, are nucleotides, and are color-balanced. Using single-length core UMIs can facilitate the use of bioinformatics tools that are built on single-length UMIs for extracting and deduplication of UMIs, regardless of variable length UMI sequences. Thus, the 4bp core sequence can serve as a recognition sequence that informs bioinformatics tools to trim 5, 6, or 7 bases (including invariant T), thereby maintaining accurate cfDNA endpoint information.The use of UMIs can enable read deduplication, single-stranded error correction, and duplex reconstruction after sequencing, thereby allowing the use of the reverse complement of the read to enhance error correction, also called double-stranded error correction. In another example, a unique dual index (UDI) is an additional sequence that can be added to a UMI-containing adaptor during library preparation to provide sample barcoding and demultiplexing of samples after sequencing. In various examples, the length of the UDI sequence is 4bp, 5bp, 6bp, 7bp, 8bp, or 12bp.

[0183] In various embodiments, the oligonucleotide adaptors may contain UMIs of 4-6 bp in length with 5' thymidine overhangs. The UMIs are designed to be non-unique (e.g., drawn from a set of specific constrained sequences).

[0184] In some embodiments, some UMIs contain one or more methylcytosine bases. The efficiency of enzymatic methylation conversion reactions (including TET oxidation and APOBEC deamination) may be evaluated based on the proportion of UMIs that do not match a specific constrained set of UMI sequences designed by UMI mismatch rate. UMI mismatch rate can be used as an embedded quality control metric to evaluate the quality of sequencing libraries. Furthermore, when perfect UMI matching is required in bioinformatics pipelines, UMI mismatch rate can be used as a filter to remove individual reads that may be of lower quality due to incomplete conversion.

[0185] In various embodiments, the UMI mismatch rate is less than 6%, less than 5%, less than 4%, less than 3%, or less than 2%.

[0186] In some embodiments, the UMI contains one or more cytosines that contain modifications that can be used to monitor enzymatic activity. Non-limiting examples of these modified bases include 5mC, 5hmC, 5fC, and 5cxmC.

[0187] In some instances, cytosines present in the adapter nucleic acid are modified with a 5-methyl group or a 5-hydroxymethyl group to prevent C to T conversion in the adapter.

[0188] In one example, cytosines present in the adapter nucleic acid are modified with 5hmC, 5gmC, 5caC, or 5cxmC groups to prevent conversion of cytosine (C) to uracil (U) in the adapter.

[0189] Figure 1A provides a generalized example of an adapter used in hydroxymethylation sequencing. The adapter may contain any of the following modified cytosines 5hmC, 5gmC, 5caC, or 5cxmC in the flow cell and primer binding regions. The cytosines in the UMI region may be unmodified or modified with 5mC, 5hmC, 5gmC, 5caC, or 5cxmC. 5m (5-methyl), 5hm (5-hydroxymethyl), 5gm (β-glucosyl-5-hydroxymethyl), 5ca (5-carboxyl), 5cxm (5-carboxymethyl), UMI (unique molecular barcode).

[0190] FIG. 1B provides an example of a process for generating adapters for hydroxymethylation sequencing. Adapters can be designed and synthesized using (i) mC nucleotides, or (ii) a combination of 5hmC, 5gmC, 5caC, or 5cxmC nucleotides at positions requiring protection from deamination. For process (i), the synthesized adapters can be oxidized and optionally ( *) can be glycosylated. For process (ii), the adapters are ready for use in ligation: C (cytosine), m (methyl), 5hm (5-hydroxymethyl), 5gm (β-glucosyl-5-hydroxymethyl or 5-(β-glucosyloxymethyl)cytosine), 5ca (5-carboxyl), 5cxm (5-carboxymethyl).

[0191] FIG. 2 provides a schematic overview of an exemplary 5hmC-seq assay. The operation of the 5hmC-seq assay begins with the adapters generated from FIG. 1B, for example, that are protected from downstream enzymatic conversion. The target enrichment operation is optional ( * ).

[0192] One advantage of this approach may be that pre-conversion adapter ligation preserves fragment endpoint and length information compared to approaches that involve bisulfite conversion followed by ssDNA adapter ligation. Considerable degradation of the nucleic acid prior to adapter ligation may result in loss of useful fragment endpoint and length information.

[0193] Enzymatic (e.g., using APOBEC) conversion of C to U may be less degradative and result in more complete and uniform coverage on sample nucleic acid fragments compared to bisulfite conversion methods. Bisulfite degradation of DNA may also not be uniform, and thus some sequences may be preferentially degraded over other sequences that contain CG dinucleotides, the very sites being interrogated in hydroxymethylation sequencing. Thus, enzymatic approaches may provide higher coverage of CpG sites than bisulfite conversion methods using the same number of unique reads, and greater uniformity of captured reads in target enrichment applications. Furthermore, non-bisulfite methods (e.g., enzymatic conversion) may provide increased resolution of biological signals, specifically, the ability to distinguish 5mC and 5hmC in nucleic acid sequences. This information and additional resolution may be beneficial in computational approaches and other methods.

[0194] In some examples, applying an enzymatic reaction to DNA or barcoded DNA that converts unmodified methylated and hydroxymethylated cytosine nucleobases of the sample DNA or barcoded DNA to uracil nucleobases includes performing an enzymatic conversion.

[0195] In various examples, glucosylation of 5hmC in nucleic acids from biological samples protects 5hmC from deamination. Deaminases can be used to convert unmodified C, 5mC, and 5hmC to U or its derivatives. Non-limiting examples of deaminases include APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like). The embodiments described herein utilize sufficient amounts of APOBEC to overcome sequence bias in deamination of unmethylated or methylated cytosine. Furthermore, embodiments involving APOBEC conversion rather than bisulfite conversion can provide substantially less damage to nucleic acids from biological samples.

[0196] In some examples, the 5hmC sequencing method includes contacting an aliquot of a nucleic acid sample with β-GT in the absence of TET dioxygenase, followed by treatment with a cytidine deaminase (e.g., APOBEC) to generate a reaction product in which substantially all 5hmC in the aliquot is glucosylated and substantially all unmodified cytosine and 5mC are converted to uracil. After PCR amplification, uracil is replaced with thymidine, and therefore cytosine and 5mC are indistinguishable when sequenced. The resulting reaction product can be sequenced and compared to a reference sequence to distinguish 5hmC from cytosine and from 5mC. Distinguishing between these moieties can allow mapping of these modified nucleotides to a reference sequence. The reference nucleic acid sequence can be obtained by sequencing a nucleic acid sample that does not react with any β-GT or deaminase. Alternatively, the reference sequence can be used for mapping when the reference sequence is a known reference nucleic acid sequence (e.g., obtained from a database of sequences or a reference genome).

[0197] B. 5hmC Nucleic Acid Sequencing Several sequencing methods can be used to identify 5hmC, including Tet-assisted bisulfite sequencing (TAB-seq), 5hmC selective chemical labeling techniques (e.g., 5hmC-seal), APOBEC-linked epigenetic sequencing (ACE-seq), and DNA immunoprecipitation-linked chemical modification-assisted bisulfite sequencing (DIP-CAB-seq). Each method may have advantages and disadvantages.

[0198] In TAB-seq, 5hmC nucleotides are protected by modification to 5-(β-glucosyloxymethyl)cytosine (5gmC) using T4 β-glucosyltransferase (β-GT), and 5mC bases are converted to 5caC using mTet1. All C and 5caC nucleotides can then be deaminated by bisulfite conversion to U or 5caU, respectively. However, since bisulfite can degrade 90-99% of DNA, while TAB-seq achieves degradation of single base 5hmC, TAB-seq may require relatively large amounts of DNA to mitigate bisulfite-mediated degradation. Thus, the high DNA mass requirement may prevent TAB-seq from being employed to sequence 5hmC in cfDNA samples, which may be analyte limiting.

[0199] In 5hmC-Seal, β-GT is used to label 5hmC with azide-modified glucose (UDP-6-N3-Glu), and the azide group allows for the subsequent covalent attachment of biotin via click chemistry. Streptavidin beads are used to affinity capture biotin-5gmC-containing DNA fragments while simultaneously washing away unbound fragments. The captured DNA fragments are then PCR amplified and sequenced. This technique does not include an operation that allows disambiguation of 5hmC from other modified / unmodified C bases using short-read sequencing methods (e.g., reading 5gmC as C). As a result, the method can only identify cfDNA fragments that contain at least one 5hmC, but the number and specific position of 5hmC are unknown. SMRT sequencing, a long-read sequencing technology, can be used to obtain single-nucleotide resolution of 5hmC from the captured DNA fragments of 5hmC-Seal. Short-read sequencing can be preferred over long-read sequencing, being more cost-effective and less prone to errors.

[0200] Similar to TAB-seq, ACE-seq employs β-GT to protect 5hmC with a glucose moiety. Unlike TAB-seq, the conversion / deamination operation in ACE-seq is mediated enzymatically by APOBEC instead of chemically mediated by bisulfite. Thus, although ACE-seq may require less input DNA than TAB-seq, the method may still have drawbacks. First, the cfDNA input amount can be very low, e.g., only about 4 μL (estimated from the difference between the total volume of the glucosylation reaction, which is about 5 μL, and the total volume of the substrate, enzyme, and enrichment buffer components, which is about 1 μL). cfDNA samples are generally in the low range of hundreds of picograms (pg) / μL (e.g., ~200 pg / μL), and thus the method may only support low cfDNA mass input (<1-2 ng) without devising a workaround to enrich the cfDNA. Therefore, this low cfDNA input amount may essentially limit the sensitivity of the method to identify very rare 5hmC in cfDNA as a biomarker in disease applications. Second, enzymatic glycosylation and deamination of cfDNA is performed before adapter ligation in ACE-seq. Generally, dsDNA-dependent adapter ligation is the first operation in NGS applications. However, if adapter ligation is performed before deamination, C in the adapter will deaminate to U, which will not be compatible with Illumina's platform sequencing application. By deaminating cfDNA before ligation, the adapter cytosine may remain unchanged. However, C to U conversion in cfDNA insert from deamination may generate a non-complementary strand. Therefore, the adapter ligation strategy after cfDNA deamination may require a non-conventional ssDNA-based ligation approach. In ACE-seq, ssDNA-based ligation can be achieved by employing the Accel Methyl-NGS kit (Swift Biosciences) to introduce Illumina adapter sequences.However, this particular ssDNA ligation method may add an unknown number of low-complexity bases to the 3' end of the ssDNA (to serve as primer binding sites for second strand synthesis), thus erasing the 3' end point information. Furthermore, requiring ssDNA-based ligation may negate the possibility of detecting the reverse complement of a given read using a double-stranded UMI strategy (because the cfDNA is denatured prior to ligation). Thus, ssDNA-based libraries may lose reverse complement information, which allows for greater sequencing error suppression.

[0201] If the test converted nucleic acid sequence is a T that corresponds to a reference C at a particular CpG locus, then the C was not methylated in the original test nucleic acid fragment. In contrast, if the test converted nucleic acid sequence and the reference sequence are both a C at a particular CpG locus, the C was hydroxymethylated in the original test nucleic acid fragment.

[0202] In some examples, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 50-500x, about 25-1000x, about 50-500x, about 250-750x, about 500-200x, about 750-1500x, or about 100-2000x. In some embodiments, the nucleic acid sequence is sequenced at a depth of more than 100x or more than 500x.

[0203] In some examples, the nucleic acid sequence of the converted nucleic acid molecule is sequenced at a depth of about 500x, about 1000x, about 2000x, about 3000x, about 4000x, about 5000x, about 6000x, about 7000x, about 8000x, about 9000x, about 10000x, or greater than 5000x.

[0204] In some examples, the nucleic acid sequence of the converted nucleic acid molecule is sequenced to a depth of about 300x unique, about 400x unique, about 500x unique, about 600x unique, about 700x unique, about 800x unique, about 900x unique, or about 1000x unique, or greater than 500x unique.

[0205] C. Hydroxymethylation Profiling In various examples, once enzymatic hydroxymethylation sequencing is completed, the assay can be used to analyze the hydroxymethylation status of nucleic acids in a biological sample. In some examples, whole genome enzymatic hydroxymethyl sequencing ("WG EHM-seq") provides high-resolution sequencing by characterizing the DNA hydroxymethylation status of nearly every cytidine nucleotide in the genome. Other targeted methods, such as targeted enzymatic hydroxymethyl sequencing ("TEHM-seq"), can be useful for methylation analysis.

[0206] The hydroxymethylation profile of cfDNA can be identified by applying sequence alignment methods to map hydroxymethyl-sequencing reads from the whole genome or targeted hydroxymethyl-sequencing to the human reference genome. Non-limiting examples of sequence alignment methods include bwa-meth, bismark, Last, GSNAP, BSMAP, NovoAlign, Bison, Metagenomic Phylogenetic Analysis (e.g., MetaPhlAn2), BLAT, Burrows-Wheeler Aligner (BWA), Bowtie, Bowtie2, Bfast, BioScope, CLC bio, Cloudburst, Eland / Eland2, GenomeMapper, GnuMap, Karma, MAQ, MOM, Mosaik, MrFAST / MrsFAST, PASS, PerM, RazerS, RMAP, SSAHA2, Segemehl, SeqMap, SHRiMP, Slider / SliderII, Srprism, Stampy, vmatch, ZOOM, and SOAP / SOAP alignment tools.

[0207] The use of double-stranded UMIs in hydroxymethyl sequencing can increase the accuracy of determining the true hydroxymethylation status of nucleic acid molecules. The method can account for errors that may be introduced during, for example, extraction (DNA damage), library preparation (end repair fill-in), enzymatic conversion (under- or over-conversion), PCR (base incorporation errors), and sequencing (base calling errors). Improving the accuracy of hydroxymethylation status determination can improve feature quantification and classifier generation for stratifying populations using these hydroxymethylation-based epigenetic sequence differences. The method does not rely on index barcodes for error correction.

[0208] D. Combination with Nucleic Acid Concentration Methods In another aspect, the method includes enrichment of the desired nucleic acid. In some embodiments, the hydroxymethyl sequencing method can be performed on a sample of nucleic acids enriched for the desired nucleic acid sequence. In some embodiments, the hydroxymethyl sequencing method includes a nucleic acid enrichment operation. In some embodiments, the nucleic acid enrichment method can be combined with a method for sequencing hydroxymethylated cell-free DNA. In some embodiments, the method includes adding affinity tags to only hydroxymethylated DNA molecules in a sample of cfDNA, enriching the DNA molecules tagged with the affinity tags, and sequencing the enriched DNA molecules. In some embodiments, the complementary nucleic acid molecules are used in enrichment methods to target genomic sequences having methylation states involved in cancer progression, detection, prognosis, or treatment response.

[0209] In some embodiments, the nucleic acid is predetermined by size, nucleobase content, or nucleic acid sequence. Certain enrichment methods can be applied in combination with methods described herein, such as U.S. Patent Application Publication US20200123616 and International Publication WO2017176630A1, each of which is incorporated herein by reference.

[0210] The terms "enrich" and "enrichment" refer to the partial purification of analytes having a certain characteristic (e.g., nucleic acids that contain hydroxymethylcytosine) from analytes that do not have that characteristic (e.g., nucleic acids that do not contain hydroxymethylcytosine).

[0211] Enrichment may increase the concentration of analytes having a feature (e.g., nucleic acids containing hydroxymethylcytosine) by at least 2-fold, at least 5-fold, or at least 10-fold compared to analytes not having the feature. After enrichment, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the analytes in the sample may have the feature used for enrichment. For example, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the nucleic acid molecules in the enriched composition may contain strands having one or more hydroxymethylcytosines modified to contain a capture tag. Other definitions of terms may appear throughout this specification.

[0212] The enrichment operation of the method can be carried out using magnetic streptavidin beads, but other supports can be used.As mentioned above, the enriched cfDNA molecules (corresponding to hydroxymethylated cfDNA molecules) can be amplified by PCR and then sequenced.In such an embodiment, the enriched cfDNA sample can be amplified using one or more primers that hybridize to the added adaptor (or its complement).In some embodiments, the enriched DNA sample is deaminated, for example, using APOBEC, before PCR amplification.This series of operations can allow base resolution determination of 5hmC modification on the enriched DNA.

[0213] In some embodiments, the deaminated condensed DNA can be amplified using one or more primers that hybridize to Y-shaped adaptors. In embodiments where Y-shaped adaptors (Y adaptors) are added, the adaptor-ligated nucleic acids can be amplified by PCR using two primers, a first primer that hybridizes to the single-stranded region of the top strand of the adaptor, and a second primer that hybridizes to the complement of the single-stranded region of the bottom strand of the Y adaptor (or hairpin adaptor after cleavage of the loop). For example, in some embodiments, the Y adaptor used can have P5 and P7 arms (sequences compatible with Illumina's sequencing platform), and the amplification products can have a P5 sequence on one side and a P7 sequence on the other. These amplification products can be hybridized to Illumina's sequencing substrate and sequenced. In some embodiments, the primer pair used for amplification can have a 3' end that hybridizes to the Y adaptor and a 5' tail that has either a P5 sequence or a P7 sequence. In these embodiments, the amplification products may also have a P5 sequence on one side and a P7 sequence on the other side. These amplification products may be hybridized to an Illumina sequencing substrate and sequenced. This amplification procedure may be performed by limited cycle PCR (e.g., 5 to 20 cycles).

[0214] Also provided is a method comprising: (a) obtaining a sample containing circulating cell-free DNA; (b) enriching hydroxymethylated DNA in the sample; and (c) independently quantifying the amount of nucleic acid in the enriched hydroxymethylated DNA that maps to (e.g., has a corresponding sequence) each of one or more target loci (e.g., at least one, at least two, at least three, at least four, at least five, or at least ten target loci). The method may further comprise (d) determining whether one or more nucleic acid sequences in the enriched hydroxymethylated DNA are over- or under-represented in the enriched hydroxymethylated DNA compared to a control. The identity of the nucleic acids over- or under-represented in the enriched hydroxymethylated DNA (and, in certain cases, the degree to which these nucleic acids are over- or under-represented in the enriched hydroxymethylated DNA) can be used to make a diagnosis, treatment decision, or prognosis. For example, in some cases, analysis of the enriched hydroxymethylated DNA may identify a signature that correlates with a phenotype, as described above. In some embodiments, the amount of nucleic acid molecules in the enriched hydroxymethylated DNA that map to each of one or more target loci (e.g., genes / intervals listed below) may be quantified by qPCR, digital PCR, arrays, sequencing, or any other quantitative method.

[0215] In some embodiments, the method may include attaching a label to DNA molecules that contain one or more hydroxymethylcytosine and methylcytosine nucleotides in a sample of cfDNA, wherein the hydroxymethylcytosine nucleotides are labeled with a first capture tag, and the methylcytosine nucleotides are labeled with a second capture tag that is different from the first capture tag to generate a labeled sample; concentrating the labeled DNA molecules; and sequencing the enriched DNA molecules. This embodiment of the method may include separately concentrating the DNA molecules that contain one or more hydroxymethylcytosine and the DNA molecules that contain one or more methylcytosine nucleotides. The label may be adapted from the above method or from Song et al. ("Simultaneous single-molecule epigenetic imaging of DNA methylation and hydroxymethylation", Proc. Natl. Acad. Sci. 2016 113:4338-43, incorporated herein by reference), and the capture tag is used instead of a fluorescent label.

[0216] In some embodiments, the enrichment method can be performed by ligating DNA to a universal adaptor, for example, an adaptor that ligates to both ends of a fragment of cfDNA. In certain cases, the universal adaptor can be performed by ligating a Y adaptor (or a hairpin adaptor) to the end of cfDNA, thereby generating a double-stranded DNA molecule with a top strand that contains a 5' tag sequence that is not the same as or complementary to the tag sequence added to the 3' end of the strand. The DNA fragments used in the initial operation of the method can be non-amplified DNA that has not been previously denatured. As shown in Figure 1A, this operation can require polishing (e.g., blunting) the ends of cfDNA with a polymerase, for example, A-tailing the fragments using Taq polymerase, and ligating T-tailed Y-adapters to the A-tailed fragments. This initial ligation operation can be performed on a limited amount of cfDNA. For example, adaptor-ligated cfDNA may contain less than 200 ng of DNA, e.g., 10 pg-200 ng, 100 pg-200 ng, 1 ng-200 ng, 5 ng-50 ng, or less than 10,000 ng (e.g., less than 5,000, less than 1,000, less than 500, less than 100, or less than 10) haploid genome equivalents, depending on the genome. In some embodiments, the method is performed using less than 50 ng of cfDNA (approximately equivalent to approximately 5 mL of plasma), or less than 10 ng of cfDNA, approximately equivalent to approximately 1 mL of plasma. For example, Newman et al. ("An ultrasensitive method For quantitating circulating tumor DNA with broad patient coverage," Nat Med. 2014 20:548-54, incorporated herein by reference) describe libraries from 7-32 ng of cfDNA isolated from 1-5 mL of plasma. This is equivalent to 2,121–9,697 haploid genomes (assuming 3.3 pg per haploid genome).The adaptors ligated onto the cfDNA may contain molecular barcodes to facilitate multiplexing and quantitative analysis of sequenced molecules. Specifically, the adaptors may be "indexed" in that they contain molecular barcodes that identify the sample to which they are ligated, allowing samples to be pooled prior to sequencing. Alternatively, or in addition, the adaptors may contain random barcodes, etc. Such adaptors may be ligated to fragments, such that substantially all fragments corresponding to a particular region are tagged with a different sequence. This allows identification of PCR duplicates and allows molecules to be counted.

[0217] In the next step of this implementation of the method, the hydroxymethylated DNA molecules in the cfDNA are labeled with a chemoselective group, for example, a group that can participate in a click reaction. This can be done by incubating the adaptor-ligated cfDNA with a DNA β-glucosyltransferase (e.g., T4 DNA β-glucosyltransferase, which is commercially available from a number of suppliers, although other DNA β-glucosyltransferases exist), and, for example, UDP-6-N3-GIU (e.g., UDP glucose containing an azide). This can be done, for example, using a protocol adapted from U.S. Patent Publication No. US20110301045, which is incorporated herein by reference, or Song et al. ("Selective chemical labeling revealing the genome-wide distribution of 5-hydroxymethylcytosine," Nat.Biotechnol.2011 29:68-72, which is incorporated herein by reference).

[0218] The next step in this implementation of the method involves adding a biotin moiety to the chemoselectively modified DNA via a cycloaddition (click) reaction. This can be done by adding a biotinylation reactant, e.g., dibenzocyclooctyne-modified biotin, directly to the glucosyltransferase reaction after the reaction is complete, e.g., after a suitable time (e.g., 30 minutes or more). In some embodiments, the biotinylation reactant can be of the general formula BLX, where B is a biotin moiety, L is a linker, and X is a group that reacts with the chemoselective group added to the cfDNA via the cycloaddition reaction. In certain cases, the linker can make the compound more soluble in an aqueous environment, and as such, can contain a polyethylene glycol (PEG) linker or equivalent. In some embodiments, the compound added can be dibenzocyclooctyne-PEG-biotin, where N is 2-10, e.g., 4. Dibenzocyclooctyne-PEG4-biotin is relatively hydrophilic and soluble in aqueous buffer up to a concentration of 0.35 mM. The compound added in this procedure does not need to contain a cleavable bond, such as a disulfide bond. In this procedure, the cycloaddition reaction can be between the azide group added to the hydroxymethylated cfDNA and the alkynyl group (e.g., dibenzocyclooctyne group) attached to the biotin moiety. This procedure can also be carried out using a protocol adapted from, for example, US Patent Publication No. US20110301045 or Song et al. ("Selective chemical labeling revealing the genome-wide distribution of 5-hydroxymethylcytosine", Nat.Biotechnol.2011 29:68-72, which is incorporated herein by reference).

[0219] The enrichment step of the method can be carried out using magnetic streptavidin beads, but other supports can be used.As described above, the enriched cfDNA molecules (corresponding to hydroxymethylated cfDNA molecules) are amplified by PCR and then sequenced.

[0220] In these embodiments, the enriched DNA sample may be amplified using one or more primers that hybridize to the added adapters (or their complements). In embodiments in which a Y adapter is added, the adapter-ligated nucleic acid may be amplified by PCR using two primers, a first primer that hybridizes to the single-stranded region of the top strand of the adapter, and a second primer that hybridizes to the complement of the single-stranded region of the bottom strand of the Y adapter (or the hairpin adapter after cleavage of the loop). For example, in some embodiments, the Y adapter used may have P5 and P7 arms (e.g., with sequences compatible with Illumina's sequencing platform), and the amplification products may have a P5 sequence on one side and a P7 sequence on the other. These amplification products may be hybridized to an Illumina sequencing substrate and sequenced. In some embodiments, the primer pair used for amplification may have a 3' end that hybridizes to the Y adapter and a 5' tail that has either a P5 sequence or a P7 sequence. In these embodiments, the amplification products may also have a P5 sequence on one side and a P7 sequence on the other side. These amplification products may be hybridized to an Illumina sequencing substrate and sequenced. This amplification procedure may be performed by limited cycle PCR (e.g., 5 to 20 cycles).

[0221] The sequencing operation can be performed using any convenient next-generation sequencing method, and can result in at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 10 million, at least 100 million, or at least 1 billion sequence reads. In some cases, the reads are paired-end reads. The primers can be used for amplification and can be compatible for use in any next-generation sequencing platform that uses primer extension, such as Illumina's reversible terminator method, Roche's pyrosequencing method (454), Life Technologies' sequencing by ligation (SOLiD platform), Life Technologies' Ion Torrent platform, or Pacific Biosciences' fluorescent base-cleavage method.Examples of such methods are described in the following references: Margulies et al. ("Genome sequencing in microfabricated high-density picolitre reactors", Nature 2005 437:376-380); Ronaghi et al. ("Real-time DNA sequencing using detection of pyrophosphate release", Anal Biochem. 1996;242:84-89); Shendure et al. ("Accurate Multiplex Polony Sequencing of an Evolved Bacterial Genome", Science 2005, 309:1728-1732); Imelfort et al. ("De novo sequencing of plant genomes using second-generation technologies", Brief Bioinform. 2009, 10:609-618); Fox et al. ("Applications of ultra-high-throughput sequencing", Methods Mol. Biol. 2009;553:79-108), Appleby et al. ("New technologies for ultra-high-throughput genotyping in plants", Methods Mol Biol. 2009;513:19-39), English et al. ("Mind the Gap: Upgrading Genomes with Pacific Biosciences RS Long-Read Sequencing Technology", PLoS ONE. 2012;7:e47768), and Morozova et al. ("Applications of next-generation sequencing technologies in functional Genomics", Genomics. 2008,92:255-264), each of which is incorporated herein by reference and may be used for an overview of the methods and specific operations of the methods, including the starting products, reagents, and final products for each of the operations.

[0222] In some embodiments, a sequenced sample may include a pool of DNA molecules from multiple samples, and the nucleic acids in the sample include molecular barcodes to indicate their source. In some embodiments, the nucleic acids may be derived from a single source (e.g., a single organism, virus, tissue, cell, subject, etc.). In other embodiments, a nucleic acid sample may be a pool of nucleic acids extracted from multiple sources (e.g., a pool of nucleic acids from multiple organisms, tissues, cells, subjects, etc.), with "plurality" meaning two or more. Thus, in some embodiments, a nucleic acid sample may contain nucleic acids from two or more sources, three or more sources, five or more sources, ten or more sources, fifty or more sources, one hundred or more sources, five hundred or more sources, one thousand or more sources, five thousand or more sources, up to about 10,000 or more sources. The molecular barcodes may allow sequences from different sources to be distinguished after they are analyzed.

[0223] The sequence reads can be analyzed by a computer, and thus, instructions for performing the operations described below can be written as a program, which can be recorded on a suitable physical computer readable storage medium.

[0224] II. Computer Systems and Machine Learning Methods A. Sample characteristics As used herein, with respect to machine learning and pattern recognition, the term "feature" may refer to an individual measurable property or characteristic of the phenomenon being observed. Features may be numerical, although structural features such as strings and graphs may be used in syntactic pattern recognition. The concept of a "feature" may be related to the concept of explanatory variables used in statistical techniques such as linear regression.

[0225] In some embodiments, the hydroxymethylation status data is characterized and processed using a trained machine learning model that is trained to classify samples into groups according to pre-specified or pre-selected biological characteristics.

[0226] In some embodiments, the set of features is identified from a nucleic acid sequence that is processed using a machine learning model. The set of features may correspond to characteristics of a nucleic acid sequence in a biological sample.

[0227] In some embodiments, the nucleic acid sequence characteristic is selected from the presence or absence or stage of cancer, or a prognosis of cancer in the individual from which the sample was obtained.

[0228] The training samples can be selected based on a desired classification, for example, as indicated by a clinical question. The different subsets can have different characteristics, for example, as determined by the labels assigned to the subsets. A first subset of training biological samples can be identified as having a particular characteristic, and a second subset of training biological samples can be identified as not having the particular characteristic. Examples of characteristics can be various diseases or disorders, but also intermediate classifications or intermediate measurements. Examples of such characteristics include, but are not limited to, the presence of cancer or the stage of cancer, or the prognosis of cancer, for example, untreated or response to a treatment for cancer. By way of example, the cancer can be colorectal cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer.

[0229] In some embodiments, the features are processed using a feature matrix for machine learning analysis.

[0230] For multiple assays, the system may identify a feature set that is processed using a machine learning model. The system may perform an assay for each molecular class and form a feature vector from the measurements. The system may process the feature vector using the machine learning model to obtain an output classification of whether the biological sample has the specified characteristic.

[0231] In some embodiments, the machine learning model outputs a classifier that distinguishes between two groups or classes of individuals or features in a population of individuals or features, in some embodiments the classifier is a trained machine learning classifier.

[0232] In some embodiments, informative loci or features of biomarkers in cancer tissues are assayed to form a profile. Receiver operating characteristic (ROC) curves can be useful to plot the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing between two populations (e.g., individuals who respond to a therapeutic agent and those who do not). Feature data across an entire population (cases and controls) can be sorted in ascending order based on the value of a single feature.

[0233] In some embodiments, the disease is advanced adenoma (AA), colorectal cancer (CRC), colorectal cancer, or inflammatory bowel disease.

[0234] The term "input features" or "feature" may refer to variables used by a model to predict an output classification (label) of a sample, e.g., disease, sequence content (e.g., mutations), proposed data collection operations, or proposed treatments. Values ​​of the variables may be determined for a sample and used to determine the classification. Examples of input features of genetic data include alignment variables related to alignment of sequence data (e.g., sequence reads) to a genome, and non-alignment variables related, e.g., to sequence content of sequence reads, protein or autoantibody measurements, or average methylation levels in a genomic region.

[0235] In various embodiments, the hydroxymethylation state in a nucleic acid sequence can be measured using any of the following: 1) a feature of a single CpG site (e.g., the ratio of 5hmC to C or % hydroxymethylation), for a CpG site, the ratio of 5hmC to 5mC, the ratio of 5hmC to total methylation (5mC+5hmC); 2) a feature of a single CH site (e.g., the ratio of 5hmC to C or % hydroxymethylation), for a CH site, the ratio of 5hmC to 5mC, the ratio of 5hmC to total methylation (5mC+5hmC); 3) a fragment-level 5hmC feature (e.g., a cfDNA fragment is called hydroxymethylated if it has ≧X 5hmC CpG sites, a cfDNA fragment is called hydroxymethylated if ≧X% of the CpG sites are 5hmC, and a fragment is called hydroxymethylated if ≧X% of the CpG sites are 5hmC). 4) Region-level 5hmC features (e.g., a cfDNA fragment is called hydroxymethylated if it has ≥X 5hmC CpG sites; a cfDNA fragment is called hydroxymethylated if ≥X% of CpG sites are 5hmC for each fragment (not just CpG sites)) 5) Region-level 5hmC features (e.g., a fragment is called hydroxymethylated if it has ≥X 5hmC CpG sites; a cfDNA fragment is called hydroxymethylated if ≥X% of CpG sites are 5hmC for each fragment) A cfDNA fragment is called hydroxymethylated if it has 5hmC sites (not just CpGs), and a cfDNA fragment is called hydroxymethylated if >= X% of Cs are 5hmC across each gene body (not just CpG sites), and for each gene body, may be quantified to include the ratio of 5hmC to C or % hydroxymethylation), the ratio of 5hmC to 5mC, the ratio of 5hmC to total methylation (5mC+5hmC), or a combination thereof, where X is any number.

[0236] In some embodiments, feature quantification across the entire gene body sequence may include exons only (e.g., by agglomerating all exons of a given gene together), transcription start site regions (e.g., a 1-kb region surrounding the TSS), enhancers, CpG shelves, CpG shores, or CpG islands.

[0237] The value of the variable can be determined for the sample and used to determine the classification.Examples of input features of genetic data include alignment variables related to the alignment of sequence data (e.g., sequence reads) to genome, and non-alignment variables related to, for example, sequence content of sequence reads, protein or autoantibody measurements, or average methylation levels in genomic regions.In various examples, genetic features such as V-plot measurements, transcription factor binding analysis, FREE-C deconvolution, cfDNA measurements across transcription start sites, and DNA hydroxymethylation levels across cfDNA fragments can be used as input features processed by machine learning methods and models.

[0238] In some examples, the sequencing information includes information regarding multiple genetic features such as, but not limited to, transcription start sites, transcription factor binding sites, chromatin open and closed states, and nucleosome positioning or occupancy.

[0239] B. Data Analysis In some embodiments, the disclosure provides a system, method, or kit with data analysis implemented in software applications, computing hardware, or both. In various embodiments, the analysis application or system comprises at least a data reception module, a data pre-processing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In some embodiments, the data reception module can comprise a computer system that connects laboratory hardware or equipment to a computer system that processes laboratory data. In some embodiments, the data pre-processing module can comprise a hardware system or computer software that performs operations on data in preparation for analysis. Examples of operations that can be applied to data in the pre-processing module include affine transformation, noise removal operations, data cleaning, reformatting, or sub-sampling. A data analysis module, which may be specialized to analyze genomic data from one or more genomic materials, can, for example, address assembled genomic sequences and perform probabilistic and statistical analyses to identify abnormal patterns associated with a disease, condition, state, risk, illness, or phenotype. The data interpretation module can use analytical methods drawn from, for example, statistics, mathematics, or biology to support understanding of relationships between identified abnormal patterns and health status, functional status, prognosis, or risk. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to create visual representations of the data that can facilitate understanding or interpretation of the results.

[0240] In various embodiments, machine learning methods are applied to distinguish between samples in a sample population, hi some embodiments, machine learning methods are applied to distinguish between healthy and advanced adenoma samples.

[0241] In some embodiments, the one or more machine learning operations used to train the methylation-based prediction engine include one or more of a generalized linear model, a generalized additive model, a non-parametric regression operation, a random forest classifier, a spatial regression operation, a Bayesian regression model, a time series analysis, a Bayesian network, a Gaussian network, a decision tree learning operation, an artificial neural network, a recurrent neural network, a reinforcement learning operation, a linear / non-linear regression operation, a support vector machine, a clustering operation, and a genetic algorithm operation.

[0242] In various embodiments, the computational method is selected from logistic regression, multiple linear regression (MLR), dimension reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier-based, wavelet, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimension reduction, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, and artificial neural networks.

[0243] In some embodiments, the methods disclosed herein can include computer analysis of nucleic acid sequencing data of samples from an individual or multiple individuals. The analysis can identify variants inferred from the sequence data to identify sequence variants based on probability modeling, statistical modeling, mechanistic models, network modeling, or statistical inference. Non-limiting examples of analysis methods include principal component analysis, autoencoder, singular value decomposition, Fourier-based, wavelet, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline variations or somatic mutations. In some embodiments, variants may refer to observed variants. Observed variants may be scientifically confirmed or reported in the literature. In some embodiments, variants may refer to putative variants associated with biological changes. Biological changes may be observed or unobserved (e.g., known or unknown). In some embodiments, putative variants may be reported in the literature but have not yet been biologically confirmed.

[0244] Alternatively, putative variants may not be reported in the literature, but can be inferred based on the computational analysis disclosed herein. In some embodiments, germline variants can refer to nucleic acids that induce natural or normal mutations.

[0245] Natural or normal variations may include, for example, skin color, hair color, and normal body weight. In some embodiments, somatic mutations may refer to nucleic acids that induce acquired or abnormal mutations. Acquired or abnormal variations may include, for example, cancer, obesity, diseases, conditions, disorders, and disorders. In some embodiments, the analysis may include distinguishing between germline variants. Germline variants may include, for example, private variants and somatic mutations. In some embodiments, the identified variants may be used by clinicians or other medical professionals to improve health care methodologies, diagnostic accuracy, and cost reduction.

[0246] Also provided herein are improved methods and computing systems or software media that can distinguish between sequence errors in nucleic acids introduced through amplification and / or sequencing technologies, somatic mutations, and germline variants. The methods provided can include simultaneously calling and scoring variants from aligned sequencing data of all samples obtained from a patient.

[0247] Samples obtained from subjects other than patients can also be used. Other samples can also be collected from subjects that have been previously analyzed by sequencing assay or targeted sequencing assay (e.g., targeted resequencing assay). The method, computing system, or software medium disclosed herein can improve the identification and accuracy of mutations or mutations (e.g., germline or somatic, including copy number mutations, single nucleotide mutations, indels, gene fusions) and the lower limit of detection by reducing the number of false positive and false negative identifications.

[0248] C. Classifier generation In some embodiments, the systems and methods provide classifiers generated based on feature information derived from methylation sequence analysis from cfDNA biological samples. The classifiers may form part of a prediction engine for distinguishing groups in a population based on methylation sequence features identified in biological samples such as cfDNA.

[0249] In some embodiments, the classifier is created by normalizing the methylation information by formatting similar portions of the methylation information into a unified format and scale, storing the normalized methylation information in a columnar database, and training a methylation prediction engine by applying one or more machine learning operations to the stored normalized methylation information, where the methylation prediction engine is created by training a methylation prediction engine that maps one or more feature combinations for a particular population, and applying the methylation prediction engine to the accessed field information to identify military-associated methylation and classify individuals into groups.

[0250] Specificity may be defined as the probability of a negative test among patients without the disease. Specificity is equal to the number of disease-free individuals who test negative divided by the total number of disease-free individuals.

[0251] In various embodiments, the model, classifier, or predictive test has a specificity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0252] Sensitivity may be defined as the probability of a positive test among those with the disease. Sensitivity is equal to the number of diseased individuals who test positive divided by the total number of diseased individuals.

[0253] In various embodiments, the model, classifier, or predictive test has a sensitivity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.

[0254] In some embodiments, the group is healthy (asymptomatic) inflammatory bowel disease, AA, or CRC.

[0255] D. Digital Processing Device In some embodiments, a digital processing device or use thereof is described herein. In some embodiments, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform the functions of the device. In some embodiments, the digital processing device may include an operating system configured to execute executable instructions. In some embodiments, the digital processing device may optionally be connected to a computer network. In some embodiments, the digital processing device may optionally be connected to the Internet so that the device accesses the World Wide Web. In some embodiments, the digital processing device may optionally be connected to a cloud computing infrastructure. In some embodiments, the digital processing device may optionally be connected to an intranet. In some embodiments, the digital processing device may optionally be connected to a data storage device.

[0256] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, handheld computers, Internet appliances, mobile smart phones, and tablet computers. Suitable tablet computers can include, for example, those having booklet, slate, and convertible configurations.

[0257] In some embodiments, the digital processing device may comprise an operating system configured to execute executable instructions. For example, the operating system may comprise software, including programs and data that manage the device's hardware and provide services for the execution of applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, NetBSD, Linux, Apple, MacOS X Server, Oracle Solaris, Windows Server, and Novell NetWare. Non-limiting examples of suitable personal computer operating systems include Microsoft Windows, Apple Mac OS X, UNIX, and UNIX-like operating systems such as GNU / Linux. In some embodiments, the operating system may be provided by cloud computing, and cloud computing resources may be provided by one or more service providers.

[0258] In some embodiments, the device may comprise a storage and / or memory device. The storage and / or memory device may be one or more physical devices used to temporarily or permanently store data or programs. In some embodiments, the device may be a volatile memory, requiring power to maintain the stored information. In some embodiments, the device may be a non-volatile memory, retaining the stored information when power is not provided to the digital processing device. In some embodiments, the non-volatile memory may include a flash memory. In some embodiments, the non-volatile memory may include a dynamic random access memory (DRAM). In some embodiments, the non-volatile memory may include a ferroelectric random access memory (FRAM). In some embodiments, the non-volatile memory may include a phase change random access memory (PRAM). In some embodiments, the device may be a storage device, including, for example, CD-ROMs, DVDs, flash memory devices, magnetic disk drives, magnetic tape drives, optical disk drives, and cloud computing-based storage. In some embodiments, the storage and / or memory device may be a combination of devices such as those disclosed herein.

[0259] In some embodiments, the digital processing device may include a display for conveying visual information to a user. In some embodiments, the display may be a cathode ray tube (CRT). In some embodiments, the display may be a liquid crystal display (LCD). In some embodiments, the display may be a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display may be an organic light emitting diode (OLED) display. In some embodiments, the OLED display may be a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In some embodiments, the display may be a plasma display. In some embodiments, the display may be a video projector. In some embodiments, the display may be a combination of devices such as those disclosed herein.

[0260] In some embodiments, the digital processing device may include an input device for receiving and processing information from a user. In some embodiments, the input device may be a keyboard. In some embodiments, the input device may be a pointing device, including, for example, a mouse, a trackball, a trackpad, a joystick, a game controller, or a stylus. In some embodiments, the input device may be a touch screen or a multi-touch screen. In some embodiments, the input device may be a microphone for capturing voice or other audio input. In some embodiments, the input device may be a video camera for capturing motion or visual input. In some embodiments, the input device may be a combination of devices such as those disclosed herein.

[0261] E. Non-transitory computer-readable recording media In some embodiments, the subject matter disclosed herein may include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of an optionally networked digital processing device. In some embodiments, the computer-readable storage medium may be a tangible component of the digital processing device. In some embodiments, the computer-readable storage medium may be optionally removable from the digital processing device. In some embodiments, the computer-readable storage medium may include, for example, CD-ROMs, DVDs, flash memory devices, solid-state memories, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In some embodiments, the programs and instructions may be encoded on the medium permanently, substantially permanently, semi-permanently, or non-permanently.

[0262] F. Computer Systems The present disclosure provides a computer system programmed to perform the methods of the present disclosure. Figure 3 shows a computer system (101) programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, or reference sequences. The computer system (101) can process various aspects of the patient data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system (101) can be a user's electronic device or a computer system located remotely relative to the electronic device. The electronic device can be a mobile electronic device.

[0263] The computer system (101) includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") (105), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (101) also includes a memory or memory location (110) (e.g., random access memory, read-only memory, flash memory), an electronic storage unit (115) (e.g., hard disk), a communication interface (120) (e.g., network adapter) for communicating with one or more other systems, and peripheral devices (125), such as cache, other memory, data storage, and / or electronic display adapters. The memory (110), storage unit (115), interface (120), and peripheral devices (125) communicate with the CPU (105) via a communication bus (solid lines), such as a motherboard. The storage unit (115) may be a data storage unit (or data repository) for storing data. The computer system (101) can be operatively coupled to a computer network ("network") (130) using the communication interface (120). The network (130) can be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The network (130) in some embodiments is a telecommunications and / or data network. The network (130) can include one or more computer servers that enable distributed computing, such as cloud computing. The network (130) can, in some embodiments, implement a peer-to-peer network with the computer system (101), which can enable devices coupled to the computer system (101) to act as clients or servers.

[0264] The CPU (105) can execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory (110). The instructions may instruct the CPU (105), which may then program or otherwise configure the CPU (105) to perform the methods of the present disclosure. Examples of operations performed by the CPU (105) may include fetch, decode, execute, and writeback.

[0265] The CPU (105) may be part of a circuit, such as an integrated circuit. One or more other components of the system (101) may be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).

[0266] The storage unit (115) can store drivers, libraries, and saved program files. The storage unit (115) can store user data, such as user preferences, and user programs. The computer system (101) can, in some embodiments, include one or more additional data storage units that are external to the computer system (101), such as located on a remote server that communicates with the computer system (101) via an intranet or the Internet.

[0267] The computer system (101) can communicate with one or more remote computer systems via the network (130). For example, the computer system (101) can communicate with a remote computer system of a user. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system (101) via the network (130).

[0268] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system (101), such as memory (110) or electronic storage unit (115). The machine executable or machine readable code can be provided in the form of software. In use, the code can be executed by the processor (105). In some embodiments, the code can be retrieved from the storage unit (115) and stored in the memory (110) for easy access by the processor (105). In some embodiments, the electronic storage unit (115) can be omitted and the machine executable instructions are stored in the memory (110).

[0269] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be interpreted or compiled during runtime. The code may be supplied in a programming language that may be selected to allow the code to be executed in a pre-compiled, interpreted, or as compiled manner.

[0270] Aspects of the systems and methods provided herein, such as the computer system (101), can be embodied in programming. Various aspects of the technology can be considered a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data transmitted or embodied on some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk.

[0271] III.How to use A. Disease Detection and Diagnosis The methods and systems provided herein may use an artificial intelligence-based approach to perform predictive analytics and analyze data obtained from a subject (patient) to generate an output of a diagnosis of the subject having cancer (e.g., CRC). For example, the application may apply a predictive algorithm to the obtained data to generate a diagnosis of the subject having cancer. The predictive algorithm may include an artificial intelligence-based predictor, such as a machine learning-based predictor, configured to process the obtained data to generate a diagnosis of the subject having cancer.

[0272] In some embodiments, the cancers detected or evaluated using the products or processes described herein include, but are not limited to, breast cancer, ovarian cancer, lung cancer, colon cancer, hyperplastic polyps, adenomas, colorectal cancer, high grade dysplasia, low grade dysplasia, prostate hyperplasia, prostate cancer, melanoma, pancreatic cancer, brain tumors (such as glioblastoma), hematological malignancies, hepatocellular carcinoma, cervical cancer, endometrial cancer, head and neck cancer, esophageal cancer, gastrointestinal stromal tumors (GIST), renal cell carcinoma (RCC), or gastric cancer. The colorectal cancer may be CRC Dukes B or Dukes CD. The hematological malignancies may be B-cell chronic lymphocytic leukemia, B-cell lymphoma-DLBCL, B-cell lymphoma-DLBCL-germinal center-like, B-cell lymphoma-DLBCL-activated B-cell-like, and Burkitt's lymphoma.

[0273] In some embodiments, the products or processes described herein may be used to detect or evaluate pre-malignant conditions such as actinic keratosis, atrophic gastritis, leukoplakia, erythroplasia, lymphomatoid granulomatosis, pre-leukemia, fibrosis, cervical dysplasia, xeroderma pigmentosum, Barrett's esophagus, colorectal polyps, or other abnormal tissue growths or lesions that are likely to develop into malignant tumors. Transforming viral infections such as HIV and HPV also present phenotypes that may be evaluated by the methods.

[0274] The cancer characterized by the methods can be, but is not limited to, a carcinoma, sarcoma, lymphoma or leukemia, germ cell tumor, blastoma, or other cancer. Carcinomas include epithelial tumors, squamous cell tumors, squamous cell carcinoma, basal cell tumors, basal cell carcinoma, transitional cell papilloma and carcinoma, adenomas and adenocarcinomas (gland), adenoma, adenocarcinoma, gastritis plastica, insulinoma, glucagonoma, gastrinoma, vipoma, cholangiocarcinoma, hepatocellular carcinoma, adenoid cystic carcinoma, adnexal carcinoid tumor, prolactinoma, oncocytoma, Hurthle cell adenoma, renal cell carcinoma, Grawitz tumor, multiple endocrine adenoma, endometrioid adenoma, adnexal and cutaneous adnexal neoplasms, mucoepidermoid neoplasms, cystic, mucinous and serous neoplasms, cystadenoma, pseudomyxoma peritonei, ductal, lobular and medullary neoplasms, acinar cell neoplasms, composite epithelial neoplasms, Warthin tumor, thymoma, specialized gonadal neoplasms These include, but are not limited to, hemophilic neoplasms, thecoma, granulosa cell tumor, arrhenoblastoma, Sertoli-Redich cell tumor, glomus tumor, paraganglioma, pheochromocytoma, glomus tumor, nevi and melanoma, melanocytic nevi, malignant melanoma, melanoma, nodular melanoma, dysplastic nevi, lentigo maligna melanoma, superficial spreading melanoma, and malignant acral lentigo melanoma. Sarcomas include, but are not limited to, Askin's tumor, botryoid tumor, chondrosarcoma, Ewing's sarcoma, malignant hemangioendothelioma, malignant schwannoma, osteosarcoma, malignant soft tissue tumor, alveolar soft tissue sarcoma, angiosarcoma, bladder sarcoma, dermatofibrosarcoma, fibroid tumor, desmoplastic small round cell tumor, epithelial sarcoma, extraskeletal chondrosarcoma, extraskeletal osteosarcoma, fibrosarcoma, hemangiopericytoma, angiosarcoma, Kaposi's sarcoma, leiomyosarcoma, liposarcoma, lymphangiomyosarcoma, lymphosarcoma, malignant fibrous tissue sarcoma, neurofibrosarcoma, and lymphofibrosarcoma.Lymphomas and leukemias include chronic lymphocytic leukemia / small lymphocytic lymphoma, B-cell prolymphocytic leukemia, lymphoplasmacytic lymphoma (e.g., Waldenstrom's macroglobulinemia), splenic marginal zone lymphoma, plasma cell myeloma, plasmacytoma, monoclonal immunoglobulin deposition disease, heavy chain disease, extranodal marginal zone B-cell lymphoma (also called MALT lymphoma), nodal marginal zone B-cell lymphoma (NMZL), follicular lymphoma, mantle cell lymphoma, diffuse large B-cell lymphoma, mediastinal (thymic) large B-cell lymphoma, intravascular large B-cell lymphoma, primary effusion lymphoma, Burkitt's lymphoma / leukemia, T-cell prolymphocytic lymphoma, and leukemia. These include, but are not limited to, mycotic leukemia, T-cell large granular lymphocytic leukemia, aggressive NK cell leukemia, adult T-cell leukemia / lymphoma, extranodal NK / T-cell lymphoma, nasal type, enteropathy type T-cell lymphoma, splenic T-cell lymphoma, blastic NK-cell lymphoma, mycosis / Sezary syndrome, primary cutaneous CD30 positive T-cell lymphoproliferative disorder, primary cutaneous anaplastic large cell lymphoma, lymphomatoid papillosis, angioimmunoblastic T-cell lymphoma, peripheral T-cell lymphoma, unspecified anaplastic large cell lymphoma, classical Hodgkin lymphoma (nodular sclerosis, mixed cytology, lymphocyte-rich, lymphocyte-depleted or non-depleted), and nodular lymphocyte-predominant Hodgkin lymphoma. Germ cell tumors include, but are not limited to, germinoma, embryonal dysplasia, seminoma, nongerminomatous germ cell tumor, embryonal carcinoma, endodermal sinus tumor, choriocarcinoma, teratoma, polygerminoma, and gonadoblastoma. Blastomas include, but are not limited to, nephroblastoma, medulloblastoma, and retinoblastoma. Other cancers include, but are not limited to, lip, laryngeal, hypopharyngeal, tongue, salivary gland, stomach, adenocarcinoma, thyroid (medullary and papillary thyroid), kidney, renal parenchymal, cervical, uterine, endometrial, choriocarcinoma, testicular, urinary tract, melanoma, brain tumors such as glioblastoma, astrocytoma, meningioma, medulloblastoma and peripheral neuroectodermal tumors, gallbladder, bronchial, multiple myeloma, basal cell tumor, teratoma, retinoblastoma, choroidal melanoma, seminoma, rhabdomyosarcoma, craniopharyngioma, osteosarcoma, chondrosarcoma, myosarcoma, liposarcoma, fibrosarcoma, Ewing's sarcoma, and plasmacytoma.

[0275] In further embodiments, the cancer under analysis may be lung cancer, including non-small cell lung cancer and small cell lung cancer (including small cell carcinoma (autocellular carcinoma), mixed small cell / large cell carcinoma, and combined small cell carcinoma), colon cancer, breast cancer, prostate cancer, liver cancer, pancreatic cancer, brain tumor, kidney cancer, ovarian cancer, stomach cancer, skin cancer, bone cancer, stomach cancer, breast cancer, pancreatic cancer, glioma, glioblastoma, hepatocellular carcinoma, papillary renal carcinoma, squamous cell carcinoma of the head and neck, leukemia, lymphoma, myeloma, or a solid tumor.

[0276] In further embodiments, the cancer is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendix cancer, astrocytoma, atypical teratoma / rhabdoid tumor, basal cell carcinoma, bladder cancer, brain stem glioma, brain tumor (including brain stem glioma, central nervous system atypical teratoma / rhabdoid tumor, central nervous system embryonal tumor, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, intermediate differentiated pineal parenchymal tumor, supraventricular primitive neuroectodermal tumor, and pineoblastoma), breast cancer, bronchial tumor, Burkitt's lymphoma, primordial neuroectodermal tumor, and pineoblastoma. Cancer of unknown origin, carcinoid tumor, cancer of unknown primary site, central nervous system atypical teratoma / rhabdoid tumor, central nervous system embryonal tumor, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine islet cell tumor, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, esthesioneuroblastoma, Ewing's sarcoma, extracranial germ cell tumor, extramaxillary germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, gastric (stomach) cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal Cellular tumors, gastrointestinal stromal tumor (GIST), gestational trophoblastic tumor, glioma, hairy cell leukemia, head and neck cancer, heart cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, pancreatic islet cell tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, malignant fibrous histiocytoma bone cancer, medulloblastoma, medulloepithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin cancer, mesothelioma, metastatic squamous cell neck cancer of unknown primary, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma / plasma cell neoplasm, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal Cavity cancer, Nasopharyngeal cancer, Neuroblastoma, Non-Hodgkin's lymphoma, Non-melanoma skin cancer, Non-small cell lung cancer, Oral cavity cancer, Oral cavity cancer, Oropharyngeal cancer, Osteosarcoma, Other brain and spinal tumors, Ovarian cancer, Ovarian epithelial cancer, Ovarian germ cell tumor, Ovarian low malignant potential tumor, Pancreatic cancer, Papillomatosis, Paranasal sinus cancer, Parathyroid cancer, Pelvic cancer, Penile cancer, Pharyngeal cancer, Moderately differentiated pineal parenchymal tumor, Pineoblastoma, Pituitary tumor, Plasma cell neoplasm / multiple myeloma, Pleuropulmonary blastoma, Primary central nervous system (CNS) lymphoma, Primary hepatocellular carcinoma, Prostate cancer, Rectal cancer, Renal cancer, Renal cell carcinoma, Respiratory tract cancer,The cancer may be retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, cervical squamous cell carcinoma, gastric (stomach) cancer, supratentorial primitive neuroectodermal tumor, T-cell lymphoma, testicular cancer, pharyngeal cancer, thymic cancer, thymoma, thyroid cancer, transitional cell carcinoma, renal pelvis and ureter transitional cell carcinoma, choriocarcinoma, ureteral cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, or Wilm's tumor. The methods of the present disclosure can be used to characterize these cancers, as well as other cancers. Thus, characterizing the phenotype can provide a diagnosis, prognosis, or theranosis of one of the cancers disclosed herein.,

[0277] The machine learning predictor can be trained using datasets from one or more sets of cohorts of patients with cancer as input, e.g., datasets generated by performing multi-analyte assays on individual biological samples, and the subject's clinical diagnosis (e.g., staging and / or tumor proportion) results as output to the machine learning predictor.

[0278] A training dataset (e.g., a dataset generated by performing a multi-analyte assay on an individual's biological sample) may be generated, for example, from one or more sets of subjects with common characteristics (features) and results (labels). The training dataset may include a set of features and labels corresponding to diagnostically relevant characteristics. The features may include features such as a particular range or category of cfDNA assay measurements that overlap or fall within each of a set of bins (genomic windows) of a reference genome, such as counts of cfDNA fragments in biological samples obtained from healthy and diseased samples. For example, a set of features collected from a given subject at a given time point may collectively serve as a diagnostic signature that may indicate the subject's identified cancer at a given time point. The features may also include labels that indicate a subject's diagnostic outcome, such as for one or more cancers.

[0279] The label may include, for example, a result, such as a clinical diagnosis (e.g., stage classification and / or tumor proportion) result of the subject. The result may include a characteristic associated with cancer in the subject. For example, the characteristic may indicate that the subject has one or more cancers.

[0280] The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). The training set may be balanced across sets of data corresponding to one or more sets of subjects (e.g., patients from different clinical sites or clinical trials). The machine learning predictor may be trained until certain predetermined conditions regarding accuracy or performance are met, such as having a minimum desired value corresponding to diagnostic accuracy measures. For example, the diagnostic accuracy measures may correspond to prediction of diagnosis, staging, or tumor fraction of one or more cancers in a subject.

[0281] Examples of diagnostic accuracy measures can include sensitivity, specificity, PPV, NPV, accuracy, and the AUC of the ROC curve, which corresponds to the diagnostic accuracy of detecting or predicting cancer (e.g., colorectal cancer).

[0282] In another aspect, the disclosure provides a method for identifying cancer in a subject, the method comprising: (a) providing a biological sample comprising cell free nucleic acid (cfNA) molecules from the subject; (b) methylation sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first set of cfNA features, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 distinct regions, each of the at least about 10 distinct regions; and (e) applying a trained algorithm to the first set of cfNA features to generate a likelihood that the subject has the cancer.

[0283] In some embodiments, the method comprises comparing a hydroxymethylation level measured in a predetermined region of interest (ROI) from a subject at risk of having a disease or cell proliferation disorder to a database of measured hydroxymethylation levels in normal or healthy subjects for a similar predetermined ROI, and determining that the subject is at increased risk of having a cell proliferation disorder by quantifying differentially hydroxymethylated nucleic acid fragments in the subject's predetermined ROI compared to the predetermined ROI of the normal or healthy subjects in the database of measured hydroxymethylation levels in normal or healthy subjects for a similar predetermined ROI.

[0284] For example, such a disease of interest may have a sensitivity for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0285] As another example, such a predetermined disease may have a specificity for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0286] As another example, such a predetermined disease may have a PPV predictive of cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0287] As another example, such a predetermined disease may have an NPV predictive of cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0288] As another example, such a predetermined disease may be one in which the AUC of a ROC curve predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) includes a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0289] In some examples of any of the aforementioned aspects, the method further includes monitoring disease progression in the subject, the monitoring being based at least in part on the gene sequence signature. In some examples, the disease is cancer.

[0290] In some embodiments, the methods described herein are useful for determining the contribution of 5-hydroxymethylation signals to total methylation signals in patient samples. Total methylation signals can be derived from various sequencing methods, including bisulfite or enzyme-based library preparation for methylation detection. The contribution of 5hmC to noise, which adversely affects diagnostic sensitivity or specificity, can be removed from the total methylation signal to improve test performance.

[0291] In some embodiments, the methods described herein are useful for 5hmC detection and can be used in a similar manner to oxidative bisulfate sequencing (oxBS-seq). Conversion of C, 5hmC, 5fC, and 5caC bases to uracil without conversion of 5mC can allow detection of only 5mC. The 5hmC signal can be subtracted from the total methylation signal to achieve a "true methyl" signal at base resolution, but with lower DNA input. Subtracting 5hmC from the total methylation signal provides a readout of the "true methyl" or 5mC signal in DNA. oxBS-seq can involve chemical oxidation of 5hmC to 5fC, followed by bisulfite conversion, which requires high DNA input.

[0292] In some embodiments, the methods described herein are useful for analyzing nucleotide resolution 5hmC alone or in combination with total methylation signals to improve gene expression prediction. Features for prediction may include 5hmC levels per CpG or fragment level and 5hmC / 5mC ratios in relevant genomic features such as promoters, enhancers, UTRs, and gene bodies.

[0293] In some embodiments, the methods described herein are useful for collecting nucleotide-level 5hmC signatures in various tissues, cell types, and cancer types, thereby enhancing the resolution of previous 5hmC tissue maps. Analysis of these data can be used for more sensitive and specific determination of tissue of origin for cancer diagnosis and prognosis.

[0294] In some embodiments, the methods described herein are useful for discovering biomarkers for patient response to cancer treatment.The abundance of 5hmC signal in cfDNA or the presence of tissue-specific 5hmC signal can be used to track residual disease after treatment for one or more cancer types.

[0295] In some embodiments, the methods described herein may use cfDNA-derived 5hmC sequence data information in drug target genes for companion diagnostics to identify patients likely to respond or respond positively to drug treatment, the efficacy of a patient's response to the drug, or patients at risk of side effects due to treatment. EXAMPLES

[0296] Example 1.5 Use of Modified Oligonucleotide Adapters to Improve Melting of hmC-Containing Nucleic Acids The methods described herein can be used to generate nucleotide-resolution 5hmC sequencing libraries from cell-free or genomic DNA molecules in patient samples. Libraries can be generated for the entire genome or for targeted regions. Analysis of 5hmC DNA modifications can have many applications, including biomarker discovery for cancer detection, tissue of origin determination, cancer prognosis, and companion diagnostic development. Characterized hydroxymethylation status data can be used as input for applications including hydroxymethylation profiling to identify characterized biomarkers of disease (including subtype stratification) or to train machine learning models useful for classifying individual samples for disease detection.

[0297] method The enzymatic hydroxymethylation sequencing (EHM-seq) method for detecting 5hmC involves the following steps: a. Enzymatic oxidation and optional glycosylation of the 5mC adapter; b. End preparation of input DNA; c. Adapter ligation to the input DNA using enzymatically oxidized adapters; d. β-glucosylation of C and 5mC in DNA molecules and protection of 5hmC by enzymatic deamination to U; and e. Sequencing of ligated DNA of transformed input may include.

[0298] A) Enzymatic oxidation of the 5mC adaptor. The enzymatic oxidation of 5mC in the adaptor involves enzymatic oxidation first to 5hmC, then to 5fC, and finally to 5caC, which may simultaneously, in the same reaction, glucosylate 5hmC to 5gmC. In this way, 5caC and 5gmC may be protected from downstream conversion to U.

[0299] Oxidation of 5mC or glucosylation to 5caC and / or 5gmC protects the adapter from downstream enzymatic conversion to U and the ligated DNA molecule can be subjected to 5hmC detection.

[0300] An alternative to enzymatically oxidizing 5mC adaptors would be to synthesize 5hmC-containing adaptors for use in subsequent adaptor ligation reactions.

[0301] B) End preparation and A-tailing of input DNA End repair uses a DNA polymerase with 3'-5' exonuclease activity to fill in the 5' overhang and remove the 3' overhang, thereby generating blunt-ended DNA. A-tailing then attaches a single A nucleotide to the 3' end, allowing for a subsequent highly efficient T / A ligation operation. Alternatively, the A-tailing operation can be omitted if blunt-end ligation is used to attach adapters to DNA molecules.

[0302] C) Adapter ligation and library preparation The enzymatically oxidized adaptors are added to the adaptor ligation reaction with the sample DNA molecules at a final concentration of 1 μM. After adaptor ligation, a cleanup is performed and the adaptor-ligated DNA molecules are eluted in a final volume.

[0303] D) Protection of 5hmC by glycosylation to 5gmC The ligated DNA is glycosylated. After glycosylation, a cleanup is performed and the glycosylated adaptor-ligated DNA molecules are eluted in the final volume.

[0304] The cleaned β-GT protected DNA is denatured followed by immediate incubation on ice. The denatured DNA is subjected to APOBEC reaction conditions to complete the enzymatic conversion.

[0305] The converted DNA can then be PCR amplified for target enrichment and / or sequencing.

[0306] Hydroxymethylation analysis / feature quantification 5hmC is preferentially expressed in genomic regions of the genome, including enhancers, promoters, and gene bodies. Useful feature quantification of the data generated by the methods described herein is used to calculate aggregate 5hmC metrics across gene bodies, such as, for example, the average hydroxymethylation level (the number of hydroxymethylated CpGs detected overlapping the gene body divided by the total number of CpGs overlapping the gene body). One possible use of this metric is to classify the disease state of a sample.

[0307] Because CpG methylation constitutes the majority of cytosine methylation in mammals, analysis of cytosine methylation and hydroxymethylation in mammalian genomes has traditionally focused on methylation of cytosines in CpG contexts. However, non-CpG methylation, i.e., CH methylation, may be biologically functional. Hydroxymethyl states in nucleic acid sequences can be feature-quantified to include average CH hydroxymethylation levels across gene bodies. Once quantified, hydroxymethylation state data can be processed for applications including hydroxymethylation profiling to identify characteristic biomarkers of disease (including subtype stratification) or to train machine learning models useful for classifying individual samples for disease detection.

Claims

1. A method for providing hydroxymethylation state data of nucleic acids in a biological sample, the method comprising: a) obtaining the biological sample containing the nucleic acid; b) ligating an oligonucleotide adapter to at least a part of the nucleic acid in the biological sample, thereby generating a ligated nucleic acid, wherein the oligonucleotide adapter comprises 5-hydroxymethylcytosine (5hmC) nucleotides, 5-(β-glucosyloxymethyl)cytosine (5gmC) nucleotides, 5-carboxycytosine (5caC) nucleotides, 5-carboxymethylcytosine (5cxmC) nucleotides, or a combination thereof; c) applying conversion conditions to convert unmethylated cytosine nucleotides and methylated cytosine nucleotides of the ligated nucleic acid to uracil nucleotides, but not converting hydroxymethylated cytosine nucleotides to uracil nucleotides, to at least a part of the ligated nucleic acid or a derivative thereof, thereby generating a converted nucleic acid; d) sequencing at least a part of the converted nucleic acid to obtain a nucleic acid sequence of the converted nucleic acid, thereby providing the hydroxymethylation state data of the nucleic acid. A method comprising the above steps.

2. The method according to claim 1, wherein the oligonucleotide adapter does not contain cytosine nucleotides in a flow cell binding region or a primer binding site of the oligonucleotide adapter.

3. The method according to claim 1, further comprising, after b) or before c), performing glucosylation of at least a part of the ligated nucleic acid with β-glucosyltransferase (β-GT) / UDP-glucose to convert the 5hmC nucleotide to the 5gmC nucleotide.

4. The method according to claim 1, wherein the conversion conditions include bisulfite treatment, enzyme treatment, or a combination thereof.

5. The method according to claim 1, wherein the oligonucleotide adapter contains the 5hmC nucleotide.

6. The method according to claim 1, wherein the oligonucleotide adapter contains the 5gmC nucleotide, the 5caC nucleotide, the 5cxmC nucleotide, or a combination thereof.

7. The method according to claim 1, wherein the conversion conditions include treatment with β-GT, cytosine dioxygenase enzyme, carboxymethyltransferase, apolipoprotein B mRNA editing catalytic polypeptide-like protein (AID / APOBEC), or a combination thereof.

8. The method according to claim 7, wherein the cytosine dioxygenase enzyme includes ten-eleven translocation protein 1 (TET1), ten-eleven translocation protein 2 (TET2), ten-eleven translocation protein 3 (TET3), or a functional variant thereof.

9. The method according to claim 1, further comprising a step of performing sequence enrichment after b) or before c).

10. The method according to claim 9, wherein the sequence enrichment includes target capture hybridization.

11. The method according to claim 1, further comprising a step of amplifying at least a part of the ligated nucleic acid before sequencing.

12. The method according to claim 1, wherein the oligonucleotide adapter is chemically synthesized using 5hmC phosphoramidite.

13. The method according to claim 1, wherein the oligonucleotide adapter includes the 5gmC nucleotide and the 5caC nucleotide, and the oligonucleotide adapter is at least partially generated by synthesizing a 5mC-containing oligonucleotide using phosphoramidite chemistry and enzymatically treating the 5mC-containing oligonucleotide with a TET enzyme and β-GT / UDP-glucose.

14. The method according to claim 1, further comprising a step of aligning the nucleic acid sequence to a reference genome.

15. The method according to claim 1, further comprising a step of training a machine learning model configured to generate hydroxymethylation state data using the hydroxymethylation state data of the nucleic acid.

16. The biological sample is obtained from or derived from a subject, and the method is processing the hydroxymethylation state data using a machine learning model configured to distinguish between a subject having a cell proliferative disorder and a subject not having the cell proliferative disorder; and detecting the cell proliferative disorder in the subject, based at least in part on the processing step, the method of claim 1

17. The method of claim 16, wherein the cell proliferative disorder includes colorectal cancer, breast cancer, ovarian cancer, prostate cancer, lung cancer, pancreatic cancer, uterine cancer, liver cancer, esophageal cancer, stomach cancer, thyroid cancer, or bladder cancer.

18. The method of claim 16, wherein the trained machine learning model is adjusted to detect the cell proliferative disorder with a preselected sensitivity and specificity.

19. The method of claim 16, further comprising treating the oligonucleotide adapter with a TET enzyme after a) or before b).

20. The method of claim 16, further comprising quantifying the hydroxymethylation state data and processing the quantified hydroxymethylation state data using a machine learning model trained to classify the biological sample based on a predetermined biological characteristic of the subject.

21. The method of claim 20, wherein the predetermined biological characteristic includes the presence or absence of precancer, the presence or absence of cancer, the stage of cancer, or the prognosis of cancer in the subject.

22. A method of generating an oligonucleotide adapter, the method comprising synthesizing an oligonucleotide comprising 5gmC nucleotides, 5caC nucleotides, 5cxmC nucleotides, or combinations thereof, at least in part by phosphoramidite chemistry, thereby generating the oligonucleotide adapter.