Methods and systems for methylation sequencing

CN122580443APending Publication Date: 2026-08-14FREENOM HLDG INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480072616.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-08
Filing Date
2024-09-13
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

[0088]本公开的另一方面提供了一种系统,所述系统包括:a)包括分类器的计算机可读介质,所述分类器用于通过使用机器学习模型基于甲基化标征组来区分患有细胞增殖性病症的对象和没有所述细胞增殖性病症的对象的群体;和b)用于执行所述计算机可读介质上存储的指令的一个或多个处理器。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122580443A_ABST
    Figure CN122580443A_ABST
Patent Text Reader

Abstract

The methods and systems presented in this paper address the limitations of nucleic acid methylation sequencing and its applications in disease detection by improving quality, sensitivity, and accuracy. The methods further include minimally destructive methylation conversion methods and nucleic acid library preparation techniques to improve methylation sequencing by minimizing signal loss and bias. Furthermore, this paper provides more accurate and complete methylation state information, thereby allowing for higher-quality feature generation for use in machine learning models and classifier generation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing This application claims the benefits of U.S. Provisional Application No. 63 / 582,948, filed September 15, 2023, and U.S. Provisional Application No. 63 / 644,275, filed May 8, 2024, each of which is incorporated herein by reference in its entirety for all purposes. Background Technology

[0002] Nucleic acid methylation can represent tumor characteristics and phenotypic status, thus holding great potential for early disease detection and / or diagnosis, as well as personalized medicine. For example, abnormal DNA methylation may be associated with various stages of cancer, from tumor initiation to progression and metastasis. Abnormal DNA methylation patterns may appear early in cancer pathogenesis, thus providing a mechanism for early cancer detection. These properties enable the use of DNA methylation patterns for cancer diagnosis. Summary of the Invention

[0003] The methods and systems presented in this paper for preparing nucleic acid libraries for methylation sequencing address the limitations of standard methylation sequencing methods by minimizing signal loss and reducing bias introduced in standard library preparation methods. These methods and systems thereby improve the quality and accuracy of nucleic acid methylation sequencing and its applications, such as disease detection. More accurate and complete information about methylation status allows for higher-quality feature generation for use in machine learning models and classifier generation.

[0004] In one aspect, this disclosure provides a method for preparing a sequencing library for methylation sequencing of one or more nucleic acid molecules of a biological sample or its derivatives, comprising: (a) Obtaining a nucleic acid composition, wherein the nucleic acid composition comprises a plurality of single-stranded nucleic acid molecules obtained or derived from the biological sample; (b) Connecting a nucleic acid adapter to one of the plurality of single-stranded nucleic acid molecules to generate an adapter-linked nucleic acid molecule, wherein the nucleic acid adapter comprises a nucleic acid resistant to base conversion via methylation enrichment; and (c) subjecting the adaptor-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed adaptor-linked nucleic acid molecule.

[0005] In some embodiments, the ligation in (b) further comprises treatment with a deoxyribonucleic acid (DNA) ligase. In some embodiments, the ligation in (b) further comprises treatment with a polynucleotide kinase.

[0006] In some embodiments, the nucleic acid adaptor comprises a double-stranded oligonucleotide including an adaptor sequence and an overhang sequence. In some embodiments, the overhang sequence is a 3' overhang sequence. In some embodiments, the overhang sequence comprises a random sequence of oligonucleotides.

[0007] In some embodiments, the nucleic acid adaptor contains one or more methylated cytosine bases. In some embodiments, the nucleic acid adaptor does not contain methylated cytosine bases or unmethylated cytosine bases.

[0008] In some embodiments, the nucleic acid adaptor includes a unique molecular identifier. In some embodiments, the unique molecular identifier is configured to enable the measurement of the enrichment efficiency of the methylation conversion method.

[0009] In some embodiments, the method further includes treating the biological sample or a derivative thereof prior to (a) to generate the plurality of single-stranded nucleic acid molecules. In some embodiments, the treatment includes denaturing double-stranded nucleic acid molecules in the biological sample or its derivative. In some embodiments, the denaturation further includes applying heat to the double-stranded nucleic acid molecules. In some embodiments, the denaturation further includes applying heat to the double-stranded nucleic acid molecules followed by rapid cooling of the denatured single-stranded nucleic acid molecules.

[0010] In some embodiments, the method further includes treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex. In some embodiments, the method does not include treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex. In some embodiments, the binding agent is a single-stranded nucleic acid binding protein (SSB).

[0011] In some embodiments, the method further includes amplifying the transformed adaptor-linked nucleic acid molecule. In some embodiments, the amplification includes polymerase chain reaction (PCR).

[0012] In some embodiments, the method further includes contacting the transformed adaptor-linked nucleic acid molecule or a derivative thereof with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence at least partially complementary to a CpG or CH locus of a reference panel. In some embodiments, the nucleic acid probe comprises an unmethylated nucleic acid probe. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at CpG loci in the CpG or CH locus of the reference panel. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at CpG loci in the CpG or CH locus of the reference panel.

[0013] In some embodiments, the method further includes determining the nucleic acid sequence of the enriched nucleic acid molecule or its derivative.

[0014] In some embodiments, the method further includes sequencing the enriched nucleic acid molecules or their derivatives to generate sequencing data. In some embodiments, the method further includes analyzing the sequencing data to generate methylation profiles of the nucleic acid molecules of the biological sample or its derivatives. In some embodiments, the methylation profiles include hypermethylation and / or hypomethylation analysis. In other embodiments, the methylation profiles include hypermethylation analysis. In still other embodiments, the methylation profiles include hypomethylation analysis. In some embodiments, the analysis further includes comparing the sequencing data with a reference sequence.

[0015] In some embodiments, the method further includes subjecting the nucleic acid molecule linked by the adaptor to an extension reaction after (b) to generate a partially double-stranded or fully double-stranded nucleic acid molecule. In some embodiments, the extension reaction is carried out in the presence of a polymerase, a variety of deoxynucleotide triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

[0016] In some embodiments, the nucleic acid molecule is deoxyribonucleic acid (DNA). In some embodiments, the DNA is cell-free DNA.

[0017] In some embodiments, the biological sample is a cell-free biological sample. In some embodiments, the cell-free biological sample is a plasma sample.

[0018] In some embodiments, the methylation enrichment method includes treatment with one or more enzymes. In some embodiments, the methylation enrichment method includes treatment with a deca-eicosyltransferase (TET) enzyme. In some embodiments, the methylation enrichment method does not include treatment with bisulfite.

[0019] In one aspect, this disclosure provides a method for preparing a sequencing library for methylation sequencing of nucleic acid molecules from biological samples or their derivatives, comprising: (a) Obtaining a nucleic acid composition, wherein the nucleic acid composition comprises a plurality of single-stranded nucleic acid molecules obtained or derived from the biological sample; (b) subjecting the single-stranded nucleic acid molecules among the plurality of single-stranded nucleic acid molecules to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating transformed single-stranded nucleic acid molecules; and (c) Linking a nucleic acid adaptor to the transformed single-stranded nucleic acid molecule to generate an adaptor-linked transformed nucleic acid molecule, wherein the nucleic acid adaptor contains a nucleic acid resistant to base transformation by the methylation enrichment method.

[0020] In some embodiments, the ligation in (c) further comprises treatment with a deoxyribonucleic acid (DNA) ligase. In some embodiments, the ligation in (c) further comprises treatment with a polynucleotide kinase.

[0021] In some embodiments, the nucleic acid adaptor comprises a double-stranded oligonucleotide including an adaptor sequence and an overhang sequence. In some embodiments, the overhang sequence is a 3' overhang sequence. In some embodiments, the overhang sequence comprises a random sequence of oligonucleotides.

[0022] In some embodiments, the nucleic acid adaptor contains one or more methylated cytosine bases. In some embodiments, the nucleic acid adaptor does not contain methylated cytosine bases or unmethylated cytosine bases.

[0023] In some embodiments, the nucleic acid adaptor includes a unique molecular identifier. In some embodiments, the unique molecular identifier is configured to enable the measurement of the enrichment efficiency of the methylation enrichment method.

[0024] In some embodiments, the method further includes treating the biological sample or a derivative thereof prior to (a) to generate the plurality of single-stranded nucleic acid molecules. In some embodiments, the treatment includes denaturing double-stranded nucleic acid molecules in the biological sample or its derivative. In some embodiments, the denaturation includes applying heat to the double-stranded nucleic acid molecules. In some embodiments, the denaturation includes applying heat to the double-stranded nucleic acid molecules to generate the plurality of single-stranded nucleic acid molecules, followed by rapid cooling of the plurality of single-stranded nucleic acid molecules.

[0025] In some embodiments, the method further includes treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex. In some embodiments, the method does not include treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex. In some embodiments, the binding agent is a single-stranded nucleic acid binding protein (SSB).

[0026] In some embodiments, the method further includes amplifying the transformed nucleic acid molecule linked by the adaptor. In some embodiments, the amplification includes polymerase chain reaction (PCR).

[0027] In some embodiments, the method further includes contacting the adapted nucleic acid molecule or its derivative linked by the adaptor with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence at least partially complementary to a CpG or CH locus of a reference group. In some embodiments, the nucleic acid probe comprises an unmethylated nucleic acid probe. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at CpG loci in the CpG or CH locus of the reference group. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at CpG loci in the CpG or CH locus of the reference group.

[0028] In some embodiments, the method further includes determining the nucleic acid sequence of the enriched nucleic acid molecule or its derivative. In some embodiments, the method further includes sequencing the enriched nucleic acid molecule or its derivative to generate sequencing data. In some embodiments, the method further includes analyzing the sequencing data to generate a methylation profile of the nucleic acid molecule in the biological sample or its derivative. In some embodiments, the methylation profile includes hypermethylation and / or hypomethylation analysis. In other embodiments, the methylation profile includes hypermethylation analysis. In still other embodiments, the methylation profile includes hypomethylation analysis. In some embodiments, the analytical method further includes comparing the sequencing data with a reference sequence.

[0029] In some embodiments, the method further includes subjecting the transformed nucleic acid molecule linked by the adaptor after (b) to an extension reaction to generate a partially double-stranded or fully double-stranded nucleic acid molecule. In some embodiments, the extension reaction is carried out in the presence of a polymerase, a variety of deoxynucleoside triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

[0030] In some embodiments, the nucleic acid molecule is deoxyribonucleic acid (DNA). In some embodiments, the DNA is cell-free DNA.

[0031] In some embodiments, the biological sample is a cell-free biological sample. In some embodiments, the cell-free biological sample is a plasma sample.

[0032] In some embodiments, the methylation enrichment method includes treatment with one or more enzymes. In some embodiments, the methylation enrichment method includes treatment with a deca-eicosyltransferase (TET) enzyme. In some embodiments, the methylation enrichment method does not include treatment with bisulfite.

[0033] In one respect, this disclosure provides a method comprising: (a) Linking a nucleic acid intransitive linker to a single-stranded nucleic acid molecule obtained or derived from a biological sample of the object to generate an intransitively linked nucleic acid molecule, wherein the nucleic acid intransitive linker comprises a nucleic acid resistant to base conversion by a methylation enrichment method; and (b) subjecting the adaptor-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed adaptor-linked nucleic acid molecule; (c) Amplify the nucleic acid molecule linked by the transformed adaptor to generate the amplified nucleic acid molecule; (d) Contact the amplified nucleic acid molecule or its derivative with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence that is at least partially complementary to the CpG or CH locus of the reference group. (e) Determine the nucleic acid sequence of the enriched nucleic acid molecule or its derivative; (f) Comparing the nucleic acid sequence of the enriched nucleic acid molecule or its derivative with a reference nucleic acid sequence; and (g) Training a machine learning model to generate a classifier that distinguishes between objects with cancer and objects without said cancer, wherein said machine learning model is trained with methylation profiles generated from: (i) a first set of nucleic acid samples from objects with said cancer, and (ii) a second set of nucleic acid samples from objects without said cancer.

[0034] In some embodiments, the methylation profile includes hypermethylation and / or hypomethylation analysis.

[0035] In other embodiments, the methylation profile includes hypermethylation analysis.

[0036] In other embodiments, the methylation profile includes hypomethylation analysis.

[0037] In some implementations, the amplification includes polymerase chain reaction (PCR).

[0038] In some embodiments, the method further includes determining the nucleic acid sequence of the enriched nucleic acid molecule or its derivative at a depth >100x.

[0039] In some embodiments, the nucleic acid probe comprises unmethylated nucleic acid. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at CpG loci in the CpG or CH loci of the reference group. In some embodiments, the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at CpG loci in the CpG or CH loci of the reference group.

[0040] In some embodiments, the method further includes sequencing the enriched nucleic acid molecules or their derivatives to generate sequencing data. In some embodiments, the method further includes analyzing the sequencing data to generate methylation profiles of the nucleic acid molecules of the biological sample or its derivatives. In some embodiments, the methylation profile includes hypermethylation and / or hypomethylation analysis. In other embodiments, the methylation profile includes hypermethylation analysis. In still other embodiments, the methylation profile includes hypomethylation analysis.

[0041] In some embodiments, the method further includes subjecting the transformed nucleic acid molecule linked by the adaptor after (a) to an extension reaction to generate a partially double-stranded or fully double-stranded nucleic acid molecule. In some embodiments, the extension reaction is carried out in the presence of a polymerase, a variety of deoxynucleoside triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

[0042] In some embodiments, the single-stranded nucleic acid molecule is deoxyribonucleic acid (DNA). In some embodiments, the DNA is cell-free DNA.

[0043] In some embodiments, the biological sample is a cell-free biological sample. In some embodiments, the cell-free biological sample is a plasma sample.

[0044] In some embodiments, the methylation enrichment method includes treatment with one or more enzymes. In some embodiments, the methylation enrichment method includes treatment with a deca-eicosyltransferase (TET) enzyme. In some embodiments, the methylation enrichment method does not include treatment with bisulfite.

[0045] In some implementations, the reference set includes CpG or CH loci associated with transcription start sites.

[0046] In some implementations, the method further includes identifying the tissue of origin of the nucleic acid molecule.

[0047] In some embodiments, the method further includes identifying the genomic location and fragment length of the nucleic acid molecule.

[0048] In some implementations, the machine learning model is trained using feature inputs selected from the following: base-wise methylation percentage of CpG, base-wise methylation percentage of CHG, base-wise methylation percentage of CHH, count or ratio of methylated CpG fragments with different counts or ratios in the observed region, conversion efficiency, low-methylated blocks, methylation level of CpG, methylation level of CHH, methylation level of CHG, fragment length, fragment midpoint, methylation level of chrM, methylation level of LINE1, methylation level of ALU, dinucleotide coverage, uniformity of coverage, global average CpG coverage, average coverage at CpG islands, CGI shelves, and CGI shores.

[0049] In some implementations, the classifier that distinguishes between subjects with the cancer and subjects without the cancer consists of a set of measurements representing methylation spectra from methylation sequencing data, which is derived from both subjects with and without the cancer. These measurements are used to generate a feature set corresponding to the characteristics of the methylation spectra. This feature set is processed by a machine learning or statistical model, which provides feature vectors that can be used as a classifier to distinguish between the groups of subjects with and without the cancer.

[0050] In some implementations, the method further includes: (a) The methylation profile of the biological sample was obtained by determination using the methylation enrichment method; (b) Classifying the methylation spectrum of the biological sample using a trained machine learning algorithm to indicate the presence of the cancer in the object; and (c) If the trained machine learning algorithm classifies the biological sample as negative for the cancer at a specified confidence level, then outputs a report identifying the biological sample as negative for the cancer.

[0051] In some embodiments, the methylation profile includes hypermethylation and / or hypomethylation analysis.

[0052] In other embodiments, the methylation profile includes hypermethylation analysis.

[0053] In other embodiments, the methylation profile includes hypomethylation analysis.

[0054] In some implementations, the method further includes: (a) Determine the baseline methylation profile of the biological sample of the object in its baseline methylation state; (b) Determine the test methylation profiles of biological samples of the object at one or more time points following the baseline methylation state; and (c) Determine the change in the test methylation spectrum compared to the baseline methylation spectrum, wherein the change indicates a change in the minor residual lesion state of the object.

[0055] In some embodiments, the methylation profile includes hypermethylation and / or hypomethylation analysis.

[0056] In other embodiments, the methylation profile includes hypermethylation analysis.

[0057] In other embodiments, the methylation profile includes hypomethylation analysis.

[0058] In some implementations, the cancer includes two or more of colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer.

[0059] In other embodiments, the cancer is colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer.

[0060] In one embodiment, the cancer is colorectal cancer. In another embodiment, the cancer is lung cancer. In yet another embodiment, the cancer is pancreatic cancer. In still another embodiment, the cancer is liver cancer.

[0061] In some implementations, the minimal residual disease status is selected from treatment response, tumor burden, postoperative tumor residue, recurrence, secondary screening, primary screening, and cancer progression.

[0062] In one aspect, this disclosure provides a method comprising: linking a nucleic acid intransitive linker to a single-stranded nucleic acid molecule obtained or derived from a biological sample of a subject to generate an intransitive linker-linked nucleic acid molecule, wherein the nucleic acid intransitive linker comprises a nucleic acid resistant to base conversion by a methylation enrichment method; subjecting the intransitive linker-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed intransitive linker-linked nucleic acid molecule; amplifying the transformed intransitive linker-linked nucleic acid molecule to generate an amplified nucleic acid molecule; and contacting the amplified nucleic acid molecule or a derivative thereof with a nucleic acid probe to generate a enriched nucleic acid molecule. The enriched nucleic acid molecules, wherein the nucleic acid probes comprise nucleic acid sequences at least partially complementary to the CpG or CH loci of the reference group; determining the nucleic acid sequences of the enriched nucleic acid molecules or their derivatives; comparing the nucleic acid sequences of the enriched nucleic acid molecules or their derivatives with the reference nucleic acid sequence; and training a machine learning model to generate a classifier that distinguishes between subjects with the indication and subjects without the indication, wherein the machine learning model is trained using methylation profiles generated from: (i) a first set of nucleic acid samples from subjects with the indication; and (ii) a second set of nucleic acid samples from subjects without the indication.

[0063] In some embodiments, the method further includes determining a methylation profile of the biological sample by the methylation enrichment method; classifying the methylation profile of the biological sample as indicating the presence of the indication in the object by a trained machine learning algorithm; and outputting a report identifying the biological sample as negative for the indication if the trained machine learning algorithm classifies the biological sample as negative for the indication at a specified confidence level.

[0064] In some embodiments, the method further includes determining a baseline methylation profile of the biological sample of the object at a baseline methylation state; determining a test methylation profile of the biological sample of the object at one or more time points after the baseline methylation state; and determining a change in the test methylation profile compared to the baseline methylation profile, wherein the change indicates a change in minimal residual disease status of the indication in the object.

[0065] In some implementations, the minimal residual disease status is selected from: treatment response, relapse, secondary screening, primary screening, and indication progression.

[0066] In some embodiments, the methylation profile includes hypermethylation analysis and / or hypomethylation analysis.

[0067] In some implementations, the indications include bowel-related diseases, immune-mediated inflammatory diseases, nervous system diseases, kidney diseases, prenatal diseases, or metabolic diseases.

[0068] In one aspect, this disclosure provides a system comprising: (a) A computer-readable medium product including the classifier of this disclosure, The classifier comprises: a set of measurements representing methylation spectra from methylation sequencing data, said methylation sequencing data being from subjects with and without said cancer; wherein said measurement set is used to generate a feature set corresponding to characteristics of the methylation spectra from subjects with and without said cancer; wherein said feature set is processed by a machine learning or statistical model, said machine learning or statistical model providing feature vectors that can be used to distinguish between subjects with and without said cancer; and (b) One or more processors for executing instructions stored on the computer-readable medium product.

[0069] In some implementations, the classifier is selected from linear discriminant analysis (LDA) classifiers, quadratic discriminant analysis (QDA) classifiers, support vector machine (SVM) classifiers, random forest (RF) classifiers, linear kernel support vector machine classifiers, first-order polynomial kernel support vector machine classifiers, second-order polynomial kernel support vector machine classifiers, ridge regression classifiers, elastic network algorithm classifiers, sequence minimum optimization algorithm classifiers, naive Bayes algorithm classifiers, and nonnegative matrix factorization (NMF) predictor algorithm classifiers.

[0070] In some implementations, the system is further configured to perform more than one of the methods described.

[0071] In some implementations, the system includes one or more processors configured to perform more than one of the methods described.

[0072] In some implementations, the system includes modules that perform the operations of any one or more of the methods described.

[0073] In one aspect, this disclosure provides a kit for detecting cancer, the kit comprising reagents for performing any of the methods described above, and instructions for detecting cancer signals.

[0074] In some implementations, the reagents are selected from primer sets, PCR reaction components, sequencing reagents, methylation enrichment reagents, and library preparation reagents.

[0075] In some implementations, the machine learning model is trained using training data obtained from: training biological samples, a first subset of the training biological samples identified as corresponding to objects suffering from a cell proliferation disorder, and a second subset of the training biological samples identified as corresponding to objects without the cell proliferation disorder.

[0076] In some implementations, the classifier is provided in a system for detecting proliferative disorders, the system comprising: a) a computer-readable medium including the classifier, the classifier being operable to classify the objects based on a methylation signature panel; and b) one or more processors for executing instructions stored on the computer-readable medium.

[0077] In some implementations, the system includes a classification circuit configured to be a machine learning classifier selected from the following: deep learning classifier, neural network classifier, linear discriminant analysis (LDA) classifier, quadratic discriminant analysis (QDA) classifier, support vector machine (SVM) classifier, random forest (RF) classifier, K-nearest neighbor, linear kernel support vector machine classifier, first-order or second-order polynomial kernel support vector machine classifier, ridge regression classifier, elastic network algorithm classifier, sequence minimum optimization algorithm classifier, Naive Bayes algorithm classifier, and principal component analysis classifier.

[0078] In some embodiments, the method further includes presenting a user's report or a graphical user interface of an electronic device. In some embodiments, the user is the object, individual, or patient.

[0079] In some implementations, the method further includes determining the likelihood of the presence or susceptibility to a proliferative disorder in the subject, individual, or patient.

[0080] In some implementations, the trained algorithm (e.g., a machine learning model or classifier) ​​includes supervised or semi-supervised machine learning algorithms. In some implementations, the supervised machine learning algorithm includes deep learning algorithms, support vector machines (SVMs), neural networks, or random forests.

[0081] In some embodiments, the method further includes providing the subject with a second diagnostic assay or procedure, such as a non-invasive assay or procedure, based at least in part on the methylation signature spectrum or the analysis described herein, including but not limited to colonoscopy, CT scan, MRI, ultrasound, or other procedures.

[0082] In some embodiments, the method further includes providing a therapeutic intervention to the subject or administering treatment to the subject, at least in part, based on the methylation signature or the analysis described herein, such as a therapeutic intervention (e.g., chemotherapy, radiotherapy, immunotherapy, or surgery) for a patient suffering from the cell proliferation disorder.

[0083] In some embodiments, the method further includes monitoring the presence or susceptibility to the proliferative disease, wherein the monitoring includes assessing the presence or susceptibility to the proliferative disease of the subject at multiple time points, wherein the assessment is based at least on the presence or susceptibility to the proliferative disease determined at each of the multiple time points.

[0084] In some implementations, the difference in the assessment of the presence or susceptibility of the subject to the proliferative disease between the plurality of time points is selected from one or more of the following clinical indications: (i) diagnosis of the presence or susceptibility of the subject to the proliferative disease, (ii) prognosis of the presence or susceptibility of the subject to the proliferative disease, and (iii) efficacy or ineffectiveness of a treatment course for the presence or susceptibility of the subject to the proliferative disease.

[0085] In some implementations, the method further includes stratifying the proliferative disease of the object by using the trained algorithm to determine the subtype of the proliferative disease of the object from multiple different subtypes or stages of the proliferative disease.

[0086] Another aspect of this disclosure provides a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements any or more of the methods described herein.

[0087] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, performs any of the methods described herein or elsewhere.

[0088] Another aspect of this disclosure provides a system comprising: a) a computer-readable medium including a classifier for distinguishing a group of objects suffering from a proliferative disorder and a group of objects without the proliferative disorder based on a methylation signature using a machine learning model; and b) one or more processors for executing instructions stored on the computer-readable medium.

[0089] Other aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of this disclosure are shown and described. As will be appreciated, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious respects, all without departing from this disclosure. Therefore, the drawings and descriptions should be considered illustrative in nature and not restrictive.

[0090] Incorporation All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated herein by reference. In the event of any conflict between a publication and patent or patent application incorporated by reference and the disclosure contained in this specification, the specification is intended to supersede and / or take precedence over any such conflicting material. Attached Figure Description

[0091] Figure 1 Schematic diagrams are provided of the example single-stranded DNA (ssDNA) library preparation method described in this article and the standard double-stranded DNA (dsDNA) library preparation method.

[0092] Figure 2 A schematic diagram of the example ssDNA library preparation workflow is provided.

[0093] Figure 3 Examples of methylation signals relative to the distance from the 3' end of a hypermethylated DNA fragment are provided when using ssDNA library preparation methods and dsDNA library preparation methods.

[0094] Figure 4 Examples comparing the captured reads relative to DNA fragment length detected using ssDNA library preparation methods and dsDNA library preparation methods are provided.

[0095] Figure 5 Examples of high methylation scores of cell-free DNA (cfDNA) extracted from samples from healthy / cancer-negative (NEG), advanced adenoma (AA), and colorectal cancer (CRC) patients, detected using the ssDNA library preparation method described herein and the dsDNA library preparation method, are provided.

[0096] Figure 6 Examples of hypermethylation scores of cfDNA extracted from healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) patients, detected using the ssDNA library preparation method and the dsDNA library preparation method described herein, are provided.

[0097] Figure 7Examples of hypomethylation scores of cfDNA extracted from samples from healthy / cancer-negative (NEG), advanced adenoma (AA), and colorectal cancer (CRC) patients, detected using the ssDNA library preparation method described herein and the dsDNA library preparation method, are provided.

[0098] Figure 8 Examples of hypomethylation scores of cfDNA extracted from samples of healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) patients, detected using the ssDNA library preparation method and the dsDNA library preparation method described herein, are provided.

[0099] Figures 9A-9B Schematic diagrams are provided illustrating two methods for converting ssDNA libraries generated using the methods disclosed herein into dsDNA libraries after ligation. Following a purification process to remove excess adaptors after adaptor ligation, single-stranded specific (SSS) primers can be annealed to the 3' end of the nucleic acid linked to the adaptor to initiate a primer extension reaction, generating a strand complementary to the nucleic acid linked to the adaptor. Figure 9A Alternatively, a non-overlapping splint adaptor can be ligated to ssDNA, followed by SSS primer annealing to the 3' end of the nucleic acid ligated to the adaptor to initiate primer extension and remove excess adaptor. Figure 9B This method allows nucleic acid amplification without initial purification procedures to remove excess adaptors.

[0100] Figure 10 Examples of protection rates of methylated cytosine (mC) in ssDNA are provided using ssDNA library workflows with second-strand synthesis (SSS) and without second-strand synthesis (SOP) after conversion of ssDNA libraries to dsDNA libraries.

[0101] Figure 11 Examples of computer systems that are programmed or otherwise configured to implement the methods provided herein are shown.

[0102] Detailed Implementation Plan This paper provides methods and systems for library preparation and sequencing of methylated regions for methylation profiling analysis of nucleic acids, such as cell-free deoxyribonucleic acid (cfDNA). These methods and systems address the limitations of existing library preparation and subsequent methylation sequencing and profiling of nucleic acids from biological samples by minimizing signal loss, reducing bias, and improving coverage, coverage uniformity, resolution, and accuracy of methylation data, thus supporting practical applications such as those described herein. The sequencing data obtained from the methods provided herein can be used for applications that classify or stratify individual populations using methylation profiling data. Such classification or stratification of individual populations may include, for example, identifying and / or detecting individuals as having a disease, staging disease progression (including detection of minimal residual disease (MRD)), or determining an individual's response to a specific treatment for a disease.

[0103] Methylation analysis can be combined with DNA sequencing to determine the likelihood that a sample is normal, tumor-derived, or disease-positive. For example, the relative abundance of methylated or unmethylated DNA fragments mapped or aligned to a specific genomic region can be used to detect or determine the likelihood of disease. Such sequencing methods can include, for example, standard DNA library preparation in which an enzymatic end-repair process is used before attaching sequencing adaptors to double-stranded DNA (dsDNA) fragments. However, these processes may fail to preserve fragment ends and may fail to capture nicked or single-stranded DNA, thereby reducing the overall methylation signal and biasing the fragment length distribution. Therefore, end repair can lead to a relative undercount of methylated molecules and an overcount of unmethylated molecules, both of which can result in an underestimation of the signal contribution from tumor-derived DNA. Consequently, the sensitivity of these methods used for cancer detection may be reduced.

[0104] The methods and systems described herein may include preparing single-stranded DNA (ssDNA) sequencing libraries and subjecting samples from the ssDNA libraries to a methylation interrogation method to convert unmethylated cytosine bases in the ssDNA to uracil bases. The methods may include denaturing double-stranded DNA (dsDNA) into ssDNA and attaching a transforming resistant nucleic acid adaptor to the ssDNA, thereby allowing evaluation of methylation in the molecule without end repair and resulting in improved and more accurate methylation analysis unaffected by the small number of methylated molecules generated from the dsDNA library preparation method. After attaching the transforming resistant nucleic acid adaptor to the ssDNA, the adaptor-ligated ssDNA may be subjected to enzymatic methylation (EM) conversion or bisulfite treatment to provide base-level resolution of DNA methylation. In one embodiment, adaptor ligation is performed prior to methylation base conversion. In another embodiment, the methylation interrogation method is performed prior to adaptor ligation. In some cases, performing adaptor ligation prior to methylation base conversion may be more preferred than performing it after methylation base conversion. Furthermore, since chemical treatments may cause more DNA damage and molecular loss than enzymatic methods, enzymatic conversion methods may be preferred over chemical conversion methods (e.g., bisulfite treatment).

[0105] For methods combining ssDNA library preparation and EM transformation, the adaptor-linked sequence can be amplified and sequenced after EM transformation. However, standard adaptors may be incompatible with these methods. Therefore, the library adaptor sequence described herein can deviate from the adaptor sequence used for standard double-stranded sequencing library preparation in terms of cytosine residue content and modification, so that the adaptor residues are resistant to transformation by chemical or enzymatic transformation methods. Furthermore, in some embodiments, the library adaptor sequence may include an additional unique molecular identifier (UMI) sequence that can be used to monitor the efficiency of the EM transformation reaction.

[0106] Various ssDNA library preparation methods can be combined with methylation interrogation methods, including, for example, adapter-based ssDNA and TdT-assisted adenosine nucleotide linker-mediated ssDNA (TACS) library preparation. Additional ssDNA library preparation methods may include single-stranded nucleic acid ligation mediated by T4-ligase using clamped oligonucleotides. These methods may include SPLinted Adapter Tagging (SPLAT) and Single-Reaction Single-Stranded Library Preparation (SRSLY) or Single-Stranded Linker Library Preparation (SALP), which utilize ssDNA-binding proteins (SSBs) and eliminate end repair for linker ligation.

[0107] Adaptase-based, TACS, and SPLAT methods for ssDNA library preparation may include additional purification steps not present in clip-and-adaptor-based ssDNA library preparation methods. These purifications can be associated with DNA loss, particularly loss of shorter DNA fragments carrying important methylation signals for disease detection. These methods may lead to further signal loss when combined with methylation analysis methods (e.g., EM conversion or bisulfite treatment prior to library preparation). Purification and recovery processes used for EM conversion and bisulfite treatment may result in greater DNA loss when processing shorter, unadaptor-ligated DNA compared to processing longer, adapter-ligated DNA. Furthermore, the compatibility of these adaptors with downstream processes (e.g., amplification and sequencing) may require several additional design considerations due to the unmethylated cytosine within the adaptor sequence transformed by EM conversion or bisulfite treatment. These additional design constraints may also apply to adaptors used for clip-ligation processes (such as SPLAT or SRSLY), which may not themselves be compatible with the enzymatic transformations downstream of library preparation. Furthermore, for example, adding SSB treatment to the SRSLY workflow may be foreign for cfDNA library preparation and may hinder the performance of the methylation conversion process.

[0108] The methods described herein may include ssDNA library preparation methods that are compatible with methylation conversion workflows without SSB treatment. These methods can be used to determine the native methylation status of cfDNA by sequencing, by minimizing molecular loss and removing biases introduced in standard library preparation methods. After denaturing dsDNA into ssDNA, the ssDNA can be ligated into adaptors designed to be insensitive to subsequent enzymatic or chemical treatments used to detect DNA methylation. These adaptor sequences may also include modules for evaluating the performance or efficiency levels of these enzymatic or chemical treatments. Following treatment, the ligated, treated DNA can be amplified and sequenced. Information obtained from this sequencing (including, but not limited to, the identity, number, size, precise ends, and relative methylation status of the DNA fragments) can be used to determine whether the cfDNA is cancer-derived.

[0109] Compared to existing library preparation methods combined with methylation conversion, the method described herein offers several advantages. Compared to existing methods, this method more faithfully preserves methylation signals from DNA extraction to library preparation, methylation conversion, and sequencing. For example, the end-repair process in dsDNA library preparation removes and inserts DNA at the 3' and 5' protrusions, respectively. These modifications result in an underestimation of true DNA methylation during sequencing of dsDNA EM-converted libraries. Since the methylation status of DNA fragments alone is sufficient to identify cancer patient-derived cfDNA, this compromised fidelity reduces the sensitivity of cancer detection methods. Similarly, reduced methylation readouts in sequenced dsDNA EM-converted libraries reduce the utility of low or no DNA methylation as a cancer detection signal. These problems are overcome by using ssDNA library preparation and EM conversion.

[0110] The method allows for the sequencing of DNA (including short ssDNA and single-stranded fragments from nicked dsDNA) that is typically lost during dsDNA library preparation. The identity and relative quantity of these fragments can serve as a tool to distinguish between DNA derived from cancer patients and DNA derived from healthy patients.

[0111] The Chinese library preparation method described in this paper preserves DNA fragment ends that may be lost or altered during the end repair process required in traditional dsDNA library preparation methods. Fragment end data can serve as a signaling source for identifying tumor-derived DNA.

[0112] Existing ssDNA library preparation methods require capturing DNA fragments after bisulfite or EM transformation to preserve adaptor sequences for downstream processing (e.g., sequencing). Because these DNA samples must undergo a purification process between transformation and adaptor ligation, they may suffer greater DNA loss compared to the methods described herein, where larger adaptor-ligated DNA molecules are washed first after ligation. This process improves molecule recovery, thereby improving the assay's capability, particularly for smaller fractions of tumor-derived signals within cell-free DNA populations. The method described herein also removes foreign reagents, such as SSB treatment, that could impair the performance of downstream assays.

[0113] Figure 1 A schematic diagram illustrating examples of dsDNA library preparation and ssDNA library preparation as described in this paper is provided. In this example, the input cfDNA can be diverse in terms of strand type (e.g., single-stranded or double-stranded), the presence of a nick, and the presence and length of the 5' and 3' overhangs.

[0114] like Figure 1As shown, the ssDNA library preparation method can preserve fragment ends and capture DNA that would not be captured using the standard dsDNA library preparation method. In the dsDNA library preparation workflow ( Figure 1 In the lower left (pictured), end-repair enzymes are used to remove or extend the ends of dsDNA fragments (pink open box) before the Y-connector is ligated to the dsDNA. These end-repair processes can lead to the exclusion of methylation signals from the ends of these fragments. In the ssDNA library preparation workflow (… Figure 1 In the lower right corner, the transformation-tolerant ssDNA compatibility adaptor is directly ligated to the ssDNA fragment. End repair is not included in the ssDNA library preparation workflow, thus preserving fragment ends and the methylation signals they generate. As described herein, the ssDNA library preparation method can therefore capture virtually all input DNA (after dsDNA denaturation), including fragments and regions of DNA lost during dsDNA end repair and library preparation (pink lines). Furthermore, the ssDNA library preparation method described herein captures native DNA methylation while excluding end repair artifacts generated by the dsDNA library preparation method.

[0115] Figure 2 A schematic diagram illustrating an example of the ssDNA library preparation workflow is provided. After nucleic acid extraction, the ssDNA is processed using an ssDNA library preparation method that utilizes a transformation-resistant clipping adaptor, followed by EM transformation, amplification, target capture, and sequencing.

[0116] Figure 3 Examples of methylation signals relative to the distance from the 3' end are provided when using ssDNA and dsDNA libraries. When CpG methylation is measured in fragments overlapping with highly methylated genomic regions, fragments in the dsDNA library exhibit a loss of methylation signal near the 3' end due to end repair (grey line (dsDNA)). In contrast, the methylation signal is uniformly high across the fragment length in the ssDNA library (green line (ssDNA)).

[0117] Figure 4 Examples comparing the captured reads relative to fragment length using ssDNA library preparation and dsDNA library preparation are provided. As shown in the figure, the ssDNA library preparation method can capture shorter cfDNA fragments corresponding to short ssDNA that were not captured by the dsDNA library preparation method.

[0118] Figure 5This paper provides a comparative example of hypermethylation model scoring of cfDNA extracted from samples of healthy / cancer-negative (NEG), advanced adenomas (AA), and colorectal cancer (CRC) patients, using the ssDNA library preparation described herein and other dsDNA library preparation methods. When processed using the ssDNA library preparation described herein and other dsDNA library preparation methods, the hypermethylation rate of cfDNA derived from selected genomic regions distinguished cfDNA derived from AA and CRC patients from cancer-negative cfDNA. This experiment demonstrates that the ssDNA library preparation method described herein produces methylation analysis (e.g., hypermethylation scoring) equivalent to other dsDNA library preparation methods. For example, when using DNA hypermethylation analysis, this workflow is equivalent to end-repair-dependent methods in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0119] Figure 6 This paper provides a comparative example of hypermethylation model scoring of cfDNA extracted from healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) patient samples, using the ssDNA library preparation method described herein and other dsDNA library preparation methods. When processed using the ssDNA library preparation method described herein and other dsDNA library preparation methods, the hypermethylation rate of cfDNA derived from selected genomic regions distinguished cfDNA derived from hepatocellular carcinoma, lung cancer, and pancreatic cancer patients from cancer-negative cfDNA. This experiment demonstrates that the ssDNA library preparation method described herein produces methylation analysis (e.g., hypermethylation scoring) equivalent to other dsDNA library preparation methods. Therefore, for example, when using DNA hypermethylation analysis, the ssDNA library preparation workflow described herein is equivalent to an end-repair-dependent method in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0120] Figure 7 This paper provides a comparative example of hypomethylation scoring of cfDNA extracted from samples from healthy / cancer-negative (NEG), advanced adenomatous (AA), and colorectal cancer (CRC) patients, using ssDNA library preparation and dsDNA library preparation as described herein. When ssDNA library preparation is used instead of dsDNA library preparation, the hypomethylation rate of cfDNA derived from selected genomic regions distinguishes AA and CRC patient-derived cfDNA from cancer-negative cfDNA. This experiment demonstrates that the ssDNA library preparation method described herein optimizes recovery, methylation analysis, quantification, and target capture. For example, when using DNA hypomethylation analysis, this workflow outperforms end-repair-dependent methods in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0121] Figure 8 Examples of hypomethylation scores for cfDNA extracted from healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) samples, detected using the ssDNA library preparation method and the dsDNA library preparation method described herein, are provided. Figure 8 As shown, when ssDNA library preparation is used instead of dsDNA library preparation, the hypomethylation rate of cfDNA derived from selected genomic regions distinguishes cfDNA derived from liver, lung, and pancreatic cancer patients from cancer-negative cfDNA. This experiment provides additional evidence that the ssDNA library preparation method described herein optimizes recovery, methylation analysis, quantification, and target capture across multiple cancer types, thereby further clarifying that this workflow can outperform end-repair-dependent methods, for example, in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients when using DNA hypomethylation analysis.

[0122] Figures 9A-9B A method is shown for converting ssDNA libraries into dsDNA libraries after ligation. The ssDNA libraries can undergo purification after ligation, after which they are added to a polymerization reaction containing a reaction buffer, nucleotides, oligonucleotide primers complementary to a portion of the 3' adaptor sequence, and polymerase. Figure 9A Alternatively, ssDNA libraries can be generated using splint adaptor sequences, allowing them to be amplified without the need for purification procedures that remove excess adaptors. Figure 9B ).

[0123] Figure 10 Examples of protection rates of methylated cytosine (mC) in ssDNA are provided using ssDNA library workflows with and without second-strand synthesis (SOP) after conversion of ssDNA libraries to dsDNA libraries. The protection rate varies with the identity of the base following methylation of the cytosine (e.g., CpA, CpG, CpC, or CpT). cfDNA_SOP represents a cell-free DNA sample using an ssDNA library workflow without second-strand synthesis; ct_SOP represents a contrived DNA sample using an ssDNA library workflow without second-strand synthesis; cfDNA_SSS represents a cell-free DNA sample using an ssDNA library with second-strand synthesis; and ct_SSS represents a contrived DNA sample using an ssDNA library with second-strand synthesis.

[0124] definition As used herein, unless the context clearly indicates otherwise, singular terms such as “a,” “an,” and “the” include both the singular and plural referents.

[0125] As used herein, the terms “cell-free plasma DNA,” “circulating cell-free DNA,” “cell-free DNA,” or “cfDNA” generally refer to DNA molecules circulating in the cell-free portion of the blood. Circulating nucleic acids in the blood can be produced from necrotic or apoptotic cells that indicate disease, such as cancer. In cancer, circulating DNA carries hallmark signs of the disease, including oncogene mutations and microsatellite alterations. This circulating DNA may be referred to as circulating tumor DNA (ctDNA). Viral genome sequences, DNA, or RNA in plasma are potential biomarkers of disease.

[0126] In some embodiments, the cell-free fraction of blood is preferably serum or plasma. As used herein, the term "cell-free fraction" of a biological sample generally refers to a substantially cell-free fraction of a biological sample. As used herein, the term "substantially cell-free" can refer to a preparation from a biological sample containing less than about 20,000 cells / mL, less than about 2,000 cells / mL, less than about 200 cells / mL, or less than about 20 cells / mL. Genomic DNA (gDNA) refers to non-fragmented DNA released from leukocytes that contaminates the cell-free fraction of blood. To mitigate gDNA contamination of samples, highly controlled sample processing workflows can be implemented, and samples can be screened for the presence of gDNA.

[0127] As used herein, the terms “detect,” “detecting,” or “detection” for condition or outcome generally include the presence of a detection indication (such as cancer), the detection status or outcome, or the tendency of the detection status or outcome.

[0128] As used herein, the term “diagnose” or “diagnosis” for a condition or outcome generally includes predicting or diagnosing a condition or outcome, determining the propensity of a condition or outcome, monitoring treatment for a patient, diagnosing a patient’s treatment response, the prognosis of a condition or outcome, its progression, and the response to a particular treatment.

[0129] As used in this article, the term "position" generally refers to the location of a nucleotide in the chain of a nucleic acid molecule that has been identified.

[0130] As used herein, the term "nucleic acid" generally refers to DNA, RNA, DNA / RNA chimeras, or hybrids that can be single-stranded (ss) or double-stranded (ds). Nucleic acids can be genomic, or derived from the genome of eukaryotic or prokaryotic cells, or synthetic, cloned, amplified, or reverse transcribed. In certain embodiments of the methods and compositions, as the context requires, nucleic acids preferably refer to genomic DNA.

[0131] As used herein, unless otherwise stated, the term "modified cytosine" generally refers to 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), formyl-modified cytosine, carboxyl-modified cytosine, 5-carboxycytosine (5caC), or cytosine modified by any other chemical group.

[0132] As used herein, the terms "methylcytosine dioxygenase," "dioxygenase," or "oxygenase" generally refer to an enzyme that converts 5mC to 5hmC. Non-limiting examples of methylcytosine dioxygenases include, for example, deca-eicosyltransfer (TET) enzymes, such as TET1, TET2, TET3, *Naegleria* TET, and their genetically engineered or non-genetically engineered variants. TET2 is an example of a methylcytosine dioxygenase that oxidizes at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, or at least 99% of all 5mC.

[0133] As used herein, the terms "methylation conversion method," "methylation enrichment method," or "methylation treatment method" generally refer to methods in which nucleic acid molecules are subjected to conditions sufficient to convert unmethylated cytosine in the nucleic acid molecule into uracil. This method can be used to distinguish between methylated and unmethylated cytosine in nucleic acid molecules.

[0134] As used herein, the terms “enzymatic methylation” or “EM conversion” or “EM-seq” generally refer to a method in which nucleic acid molecules are subjected to conditions sufficient to convert unmethylated cytosine in the nucleic acid molecule to uracil by treatment with one or more enzymes (e.g., TET enzymes). In some cases, the method does not include treatment with bisulfite (e.g., chemical treatment).

[0135] As used herein, the terms “transformation-resistant adaptor,” “transformation-resistant primer,” “transformation-tolerant adaptor,” or “transformation-tolerant primer” generally refer to nucleic acid molecules used as adaptors or primers, respectively. These nucleic acid molecules are resistant to methylation transformation or alteration by methylation enrichment methods (e.g., “deamination-resistant modified cytosine” or methylation-resistant nucleotides).

[0136] As used herein, the term "deamination-resistant modified cytosine" refers to one or more modified cytosine nucleotides in a transformation-resistant adaptor whose base pairing specificity is altered chemically or enzymatically without treatment with a methylating agent. As examples of non-limiting examples, propynyl-C and pyrrolo-C are deamination-resistant modified cytosines that are transformation-resistant nucleotides and can be included in transformation-resistant adaptors used in the library preparation methods disclosed herein. These deamination-resistant modified cytosines within the transformation-resistant adaptor are not deamination upon exposure to sodium bisulfite or a methylating invertase (e.g., an APOBEC-like enzyme).

[0137] As used herein, the term "methylation-resistant nucleotide" generally refers to a nucleotide containing a base-pairing specificity of nucleic acid bases that is not chemically altered by treatment with a methylation-resistant agent. Methylation-resistant nucleotides can be incorporated into primer extension reactions via nick translation enzymes. For example, 5-methylcytosine (5mC) is a transformation-resistant nucleotide that can be used in conjunction with sodium bisulfite or enzymatic methylation conversion. Therefore, 5-methylcytosine is not deaminated upon exposure to sodium bisulfite or methylation-resistant enzymes. As an alternative to incorporating modified nucleotide bases to prevent base conversion, transformation-resistant adaptors or transformation-resistant primers incorporate only unmodified bases to allow complete base conversion during the transformation reaction in methylation sequencing. "Unmodified bases" in the adaptor / primer DNA sequence refer to conventional guanine, adenine, cytosine, and thymine.

[0138] As used herein, the term "cytidine deaminase" generally refers to an enzyme that deaminates cytosine (C) to form uracil (U). Non-limiting examples of cytidine deaminases include cytidine deaminases of the apolipoprotein B mRNA editing enzyme catalyzing polypeptide (APOBEC) family, such as APOBEC3A. In any embodiment, the cytidine deaminase described herein may have an amino acid sequence that is at least 90% identical (e.g., at least 95% identical) to the amino acid sequence of GenBank accession number AKE33285.1 (which is the sequence of human APOBEC3A). In some embodiments, the cytidine deaminase described herein converts unmodified cytosine to uracil with an efficiency of at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% (preferably at least 99%).

[0139] As used herein, the term "glucosyltransferase" or "GT" generally refers to an enzyme that catalyzes the transfer of β-D-glucosyl or α-D-glucosyl residues from UDP-glucose to 5hmC residues to form 5ghmC. APOBEC can convert 5hmC to U at a low rate relative to the conversion of C or 5mC to U. An example of GT is T4-betaGT (βGT). In one instance, GT can be used concurrently with a dioxygenase. This combination ensures that deamination of 5hmC is blocked, such that less than 5%, less than 3%, or less than 1% of 5hmC is converted to U by the deaminase. In another instance, GT can be used together with a dioxygenase in the same reaction mixture with DNA, so that the dioxygenase converts 5mC to 5hmC and 5caC, while GT converts any remaining 5hmC to 5ghmC, ensuring that only cytosine is deaminated.

[0140] As used in this article, the terms “a portion” and “an equal part” of a nucleic acid sample are generally intended to have the same meaning and can be used interchangeably.

[0141] As used in this article, the term "comparison" generally refers to analyzing two or more sequences relative to each other. In some cases, comparison can be performed by aligning two or more sequences together, such that the nucleotides at corresponding positions are compared.

[0142] As used herein, the term "reference sequence" generally refers to the sequence of the fragment being analyzed. Reference sequences can be obtained from public databases or can be sequenced separately as part of an experiment. In some cases, the reference sequence can be hypothetical so that it can be computationally deaminated (e.g., changing C to U or T, etc.) to allow for sequence alignment.

[0143] As used herein, the terms “G,” “A,” “T,” “U,” “C,” “5mC,” “5fC,” “5caC,” “5hmC,” and “5ghmC” generally refer to nucleotides containing guanidine (G), adenine (A), thymine (T), uracil (U), cytosine (C), 5-methylcytosine, 5-formylcytosine, 5-carboxycytosine (5caC), 5-hydroxymethylcytosine, and 5-glucosylhydroxymethylcytosine, respectively. For clarity, C, 5fC, 5caC, 5mC, and 5ghmC are each distinct parts.

[0144] As used in this article, the term "minimal residual disease" or "MRD" generally refers to a small number of cancer cells remaining in the body after cancer treatment. MRD testing can be performed to determine the effectiveness of cancer treatment and guide further treatment plans. Various indicators can be used to assess MRD, including but not limited to treatment response, tumor burden, postoperative tumor residue, recurrence, secondary screening, primary screening, and cancer progression.

[0145] As used in this article, the terms “next-generation sequencing” or “NGS” are generally used for sequencing libraries of genomic fragments smaller than 1 kb in size.

[0146] As used herein, the terms “healthy” or “normal” generally refer to an object or sample derived from a disease that is free from it. While health is a dynamic state, the term can refer to a pathological state in which an object lacks the referred disease state (e.g., cancer). In one instance, when referring to the methylation profile of an object classifying an object with cancer, the term “healthy” refers to an individual lacking cancer (such as CRC). While other diseases or health states may be present in the object, the term “healthy” can indicate the absence of a stated disease in order to provide a comparison or classification between objects with and without disease states and samples derived from them.

[0147] As used herein, the term "threshold" generally refers to a value selected to distinguish, separate, or differentiate between two populations. In some embodiments, a threshold distinguishes methylation status between disease (e.g., malignant) states and non-disease (e.g., healthy) states. In some embodiments, a threshold distinguishes disease stages (e.g., stage 1, stage 2, stage 3, or stage 4). Thresholds can be set according to the disease involved and can be determined computationally based on early analyses, such as early analyses of the training set, or based on an input set with known characteristics (e.g., healthy, disease, or disease stage). Thresholds can also be set for gene regions based on predicted methylation values ​​at specific sites. The threshold can be different for each methylation site, and data from multiple sites can be combined in the final analysis.

[0148] While any methods and materials similar to or equivalent to those described herein may be used to practice or test this teaching, this article describes some example methods and materials.

[0149] References to any publication refer to its disclosure prior to the filing date and should not be construed as an admission that this claim does not claim rights prior to such publications due to prior inventions. Furthermore, the provided publication date may differ from the actual publication date, which can be independently verified.

[0150] As will be apparent to those skilled in the art upon reading this disclosure, each individual embodiment described and illustrated herein has its own components and features, which can be readily separated from or combined with features of any other several embodiments without departing from the scope or spirit of this teaching. Any method described may be performed in the order of the events described or in any other logically possible order.

[0151] Cell-free DNA (cfDNA) sequencing can be a useful tool for cancer detection. DNA methylation analysis can be combined with sequencing to determine whether a portion of cfDNA may be precancerous or tumor-derived. Many standard DNA library preparation methods use enzymatic end-repair processes to attach sequencing adaptors. However, these processes may remove native methylation signals from DNA fragments, fail to faithfully preserve molecular ends, and cannot capture nicked or single-stranded DNA. These factors reduce the sensitivity of these methods for cancer detection. Figure 1 ).

[0152] DNA methylation is a covalent modification of DNA and a stable genetic marker that plays a crucial role in suppressing gene expression and regulating chromatin architecture. In humans, DNA methylation primarily occurs at cytosine residues in CpG dinucleotides. Unlike other dinucleotides, CpGs are unevenly distributed throughout the genome and can be concentrated in short, CpG-rich DNA regions called CpG islands. Typically, most CpG sites in the genome are approximately 70-75% methylated. However, methylation patterns vary by cell type, reflecting their roles in regulating cell type-specific gene expression. In this way, the methylome of a cell can program the terminal differentiation state of the cell into, for example, neurons, muscle cells, immune cells, etc.

[0153] Furthermore, various cell subtypes within a tissue can exhibit different methylation patterns. In cancer cells, CpG methylation may be dysregulated, and abnormal methylation patterns are among the earliest events in tumorigenesis. The methylation profile of a given cancer type most closely resembles that of the tissue of origin. Therefore, abnormal methylation markers on cfDNA fragments can be used to differentiate cancer cells from normal cells and determine the tissue of origin. Typically, global CpG methylation levels are reduced in cancer cells, but at specific loci, the average methylation level (or % methylation) at a particular CpG site in cancer cells may vary relative to a matching normal cell. Profiling of differentially methylated CpGs (DMC; single site) or differentially methylated regions (DMR; more than one site in a local region) between normal and diseased cells allows for the identification of biomarkers for disease. This approach has led to the development of the SEPT9 gene methylation assay (Epi proColon), the first FDA-approved blood-based diagnostic method for colorectal cancer (CRC).

[0154] As used herein, a methylation profile may include both and / or independently include hypomethylation and hypermethylation from DNA analysis. Hypomethylation and hypermethylation are relative terms and indicate less (low) or more (high) methylation compared to a reference standard DNA. Specifically applied to cancer epigenetics, this reference standard may be DNA isolated from normal (e.g., non-cancerous) tissue.

[0155] Bisulfite conversion or bisulfite sequencing can be used for DNA methylation analysis. Bisulfite sequencing is a convenient and efficient method for mapping DNA methylation to individual bases. Unfortunately, bisulfite conversion is a harsh and destructive method for cfDNA, resulting in >90% DNA degradation in the sample. The two main methods for constructing bisulfite sequencing libraries are: (1) bisulfite conversion of DNA prior to library construction, which requires the construction of ssDNA libraries; and (2) bisulfite conversion of DNA after dsDNA adaptor ligation. Both cases involve severe DNA degradation, which can be particularly problematic for cfDNA, which is present in very low concentrations in plasma and is a limited resource in liquid biopsy applications.

[0156] Alternatively, enzymatic methylation (EM) conversion can be used for DNA methylation analysis and sequencing. In one embodiment, methylation conversion is mediated by a non-destructive enzymatic reaction, for example, using a deca-eicosine translocation (TET) enzyme and a cytosine deaminase (e.g., a cytidine deaminase from the apolipoprotein B mRNA editing peptide (APOBEC) family) to convert unmethylated (non-methylated) cytosine to uracil. Other embodiments, such as TET-assisted pyridineborane sequencing (TAPS), combine enzymatic reactions (e.g., TET treatment) with chemical treatments (e.g., the use of pyridineborane).

[0157] The advent of next-generation DNA sequencing has provided advances in clinical medicine and basic research. However, while this technology can generate hundreds of billions of nucleotides in a single experiment, an error rate of approximately 1% results in hundreds of millions of sequencing errors. In some applications, such errors are tolerable, but they become a major problem for “deep sequencing” of genetically heterogeneous mixtures, such as tumors or mixed microbial communities.

[0158] Analyzing variants and methylation status in cfDNA using existing methods requires two separate sequencing assays and two separate cfDNA pools. This can be prohibitively expensive in terms of plasma / cfDNA input and associated costs. Furthermore, DNA degradation via bisulfite treatment may reduce the sensitivity of variant-calling methods that can manipulate bisulfite-converted DNA sequencing data (as opposed to enzymatic conversion). Therefore, improved methods for analyzing cfDNA methylation are needed to preserve the integrity of sample nucleic acids and achieve improved accuracy in methylation status analysis at the whole-genome or targeted levels.

[0159] As used herein, the terms “comprising,” “including,” “having,” “containing,” “comprising,” or variations thereof are used in the detailed description and / or claims, and such terms are intended to be interpreted as inclusive terms in a manner similar to the term “comprising.”

[0160] When a range is used in this disclosure, the range may be expressed herein as from “about” a particular value and / or to “about” another particular value. When such a range is expressed in this manner, another embodiment includes from said one particular value and / or to said other particular value. In embodiments where values ​​are expressed as approximate values, for example by using the antecedent “about,” it will be understood that the specific value forms another embodiment. It will be further understood that each endpoint of the range is significant relative to and independent of the other endpoint.

[0161] Several aspects of the ssDNA library preparation methods and systems disclosed herein have been described above with reference to illustrative examples. It should be understood that numerous specific details, relationships, and methods have been set forth to provide a comprehensive understanding of the disclosed methods and systems. However, those skilled in the art will readily recognize that the methods and systems can be practiced without one or more of these specific details or using other methods. This disclosure is not limited to a specific order of actions or operations, as some actions may occur in a different order and / or simultaneously with other actions or operations. Furthermore, not all specified actions or operations are necessary to perform the methods described according to this disclosure.

[0162] I. Library preparation for enzyme-mediated methylation sequencing In the first aspect, methods for preparing sequencing libraries are provided. The methods described herein provide ssDNA libraries acceptable for both next-generation unmethylated and methylated sequencing applications, thereby providing sequencing data for both applications from a single sample. The resulting raw sequencing data can be used for methylation status analysis and existing cfDNA analyses, such as copy number alteration, germline variant detection, somatic variant detection, nucleosome localization, transcription factor profiling, chromatin immunoprecipitation, etc.

[0163] Connector linkers for targeted sequencing applications On the one hand, this method preserves the integrity and information of the nucleic acid sequences used for methylation profiling analysis. In one example, combining ssDNA adaptor ligation before enzymatic transformation preserves fragment endpoint information while increasing library complexity for target enrichment (or direct whole-genome sequencing), thereby providing higher sensitivity to detect rare events such as methylated ctDNA. The advantages of ssDNA adaptor ligation and a comparison with adaptor ligation before methylation transformation are shown in [the table / example]. Figure 1 middle.

[0164] In one instance, a nucleic acid adaptor is ligated to the 5' and 3' ends of a population of nucleic acid fragments in a biological sample to produce a sequencing library. In another instance, a set of nucleic acid adaptors is ligated to nucleic acid fragments in the sample. The nucleic acid adaptors described herein may optionally be transformation-resistant nucleic acid adaptors. In other instances, the nucleic acid adaptors may optionally be transformation-resistant adaptors containing one or more deamination-resistant modified cytosines, including but not limited to propynyl-C and pyrrolo-C. The set of adaptors may include 4-base-pair (bp), 5-bp, and 6-bp unique molecular identifier (UMI) sequences of equal portions. The UMI may be located at a neighboring nucleic acid fragment inserted into the library. During sequencing, the UMI is also sequenced as part of the 5' end read. The set of adaptors may include a single-length core UMI, which can reduce sequencing complexity at positions corresponding to non-variant thymidines that lead to reduced sequencing quality. The first 4 bp of each UMI together may include a 4-bp (e.g., single-length) core UMI sequence set with an edit distance greater than or equal to 2 and which is nucleotide and color balanced. Using a single-length core UMI in the presence of variable-length UMI sequences can facilitate UMI extraction and deduplication using bioinformatics tools built for single-length UMIs. Therefore, a 4-bp core sequence can serve as a recognition sequence, instructing bioinformatics tools to trim the 5, 6, or 7 bp sequence, thus preserving accurate cfDNA endpoint information. A schematic diagram illustrating the interleaved adaptor is shown below. Figure 2 The use of UMIs allows for read duplication and single-strand error correction. In another instance, a unique dual index (UDI) is an additional sequence that can be added during library preparation to the adaptor containing the UMI to provide sample barcoding and demultiplexing after sequencing. In various instances, the UDI sequence length is greater than or equal to 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, or 12 bp. In various instances, the UDI sequence length is less than or equal to 12 bp, 8 bp, 7 bp, 6 bp, 5 bp, or 4 bp.

[0165] In various implementations, the nucleic acid adaptor may include a UMI of 4 to 6 bp in length. The UMI may be designed to be non-unique (e.g., extracted from a specific, constrained set of sequences).

[0166] In one implementation, some UMIs contain one or more methylcytosine bases. The efficiency of the enzymatic methylation conversion reaction (including TET oxidation and APOBEC deamination) can be assessed by the UMI mismatch rate, which is the percentage of UMIs that do not match a specific, constrained, designed UMI sequence set. The UMI mismatch rate can be used as an embedded quality control metric to assess sequencing library quality. Furthermore, if perfect UMI matching is required in the bioinformatics pipeline, the UMI mismatch rate can be used as a filter to remove individual reads that may be of lower quality due to incomplete conversion.

[0167] In various implementation schemes, the UMI mismatch rate is less than or equal to 6%, less than or equal to 5%, less than or equal to 4%, less than or equal to 3%, or less than or equal to 2%.

[0168] In another embodiment, the UMI contains one or more modified cytosine bases that can be used to monitor enzymatic activity. Non-limiting examples of such modified bases include 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxycytosine.

[0169] The efficiency of subsequent processes (such as DNA purification or enzymatic methylation) can be improved by converting the ssDNA library into a dsDNA library. In some embodiments, the conversion of ssDNA to dsDNA can be performed using DNA polymerase and primers complementary to the 3' end of the adaptor sequence. The primers anneal at the 3' end of the adaptor sequence to the nucleic acid linked by the adaptor to initiate an extension reaction in the presence of polymerase and deoxynucleoside triphosphates (dNTPs). In some embodiments, the conversion of ssDNA to dsDNA can occur after a purification (“cleaning”) operation between ligation and DNA polymerization to remove free, excess adaptors and displace the reaction buffer. In other embodiments, the ssDNA to dsDNA conversion reaction components (including dNTPs, polymerase, and primers) are added directly to the ligation mixture without a purification operation. In some embodiments, the 3' adaptor splint component is truncated such that the splint component does not completely overlap with the primer annealing site of the primer extension. In this way, primers complementary to the 3' end of the adaptor sequence can anneal at the primer annealing site. Second-strand synthesis (SSS) primers may or may not have end modifiers at the 3', 5', or both ends.

[0170] II. Targeted methylation sequencing In targeted methylation sequencing methods, a target region in a biological sample (e.g., cfDNA) is analyzed to determine the methylation status of the target gene sequence. In some implementations, the target region comprises consecutive nucleotides of the target region of interest or hybridizes to consecutive nucleotides of the target region of interest under stringent conditions, such as at least about 16 consecutive nucleotides of the target region of interest. In various instances, targeted sequencing can be performed using hybridization capture and amplicon sequencing methods.

[0171] A. Hybrid capture The hybridization methods described herein can be used for various forms of nucleic acid hybridization, such as in-solution hybridization and hybridization on solid supports (e.g., Northern, Southern, and in situ hybridization on membranes, microarrays, and cell / tissue slides). Specifically, this method is suitable for in-solution hybrid capture targeting enriched genomic DNA sequences (e.g., exons) used in targeted next-generation sequencing. For hybrid capture methods, cell-free nucleic acid samples undergo library preparation. As used herein, “library preparation” includes adaptor ligation of ssDNA as described herein, or any other preparation of cell-free DNA to allow for subsequent sequencing of the DNA. In some instances, the prepared cell-free nucleic acid library sequences contain adaptors, sequence tags, and index barcodes linked to the cell-free nucleic acid sample molecules. Various commercially available kits are available to facilitate library preparation for next-generation sequencing methods. Next-generation sequencing library construction may include the preparation of nucleic acid targets using a series of coordinated enzymatic reactions to produce random sets of DNA fragments of specific sizes for high-throughput sequencing. Advances and developments in various library preparation technologies have expanded the applications of next-generation sequencing to fields such as transcriptomics and epigenetics.

[0172] Advances in sequencing technology have led to changes and improvements in library preparation. The next-generation sequencing library preparation kits used in this article include those developed by companies such as Agilent, Bioo Scientific, Claret Bioscience, Kapa Biosystems, New England Biolabs, Illumina, Life Technologies, Pacific Biosciences, and Roche.

[0173] In various instances of targeted gene capture groups, a variety of library preparation kits can be selected from Nextera Flex (Illumina), IonAmpliseq (Thermo Fisher Scientific) and Genexus (Thermo Fisher Scientific), Agilent ClearSeq (Illumina), Agilent SureSelect Capture (Illumina), Archer FusionPlex (Illumina), BiooScientific NEXTflex (Illumina), IDT xGen (Illumina), Illumina TruSight (Illumina), Nimblegene SeqCap (Illumina) and Qiagen GeneRead (Illumina).

[0174] In some embodiments, hybridization capture methods employ specific probes against prepared library sequences. As used herein, the term "specific probe" can refer to a probe specific to known methylation sites. In some embodiments, specific probes are designed based on using the human genome as a reference sequence and a designated genomic region known to have methylation sites as a target sequence. Specifically, a genomic region known to have methylation sites may include at least one of the following: promoter regions, CpG island regions, CGI shore regions, and imprinted gene regions. Therefore, when hybridization capture is performed using specific probes of some embodiments, sequences in the sample genome complementary to the target sequence (e.g., regions in the sample genome known to have methylation sites (also referred to herein as "designated genomic regions")) can be efficiently captured.

[0175] As an example, the methylated regions described herein can be used to design specific probes. In some embodiments, specific probes are designed using commercially available methods (e.g., the eArray system). The probe length is sufficient to hybridize with the methylated region of interest with adequate specificity. In various examples, the probe is a decameric, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, or 20-mer. In other embodiments, the probe length is 50-200 nucleotides. In some embodiments, the probe length is 80-150 nucleotides. In still other embodiments, the probe length is 100-130 nucleotides. In some embodiments, the probe length is greater than or equal to 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides. In some implementations, the probe length is less than or equal to 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, or 50 nucleotides.

[0176] Target regions for methylation analysis can be screened using database resources such as gene ontology. Based on the principle of complementary base pairing, single-stranded capture probes can be complementary to single-stranded target sequences to successfully capture the target region. In some embodiments, the designed probes can be configured as solid-state capture chips (where the probes are fixed to a solid support) or liquid-state capture chips (where the probes are free in a liquid). However, due to limiting factors such as probe length, probe density, and high cost, solid-state capture chips are rarely used, while liquid-state capture chips are used more frequently.

[0177] In some implementations, GC-rich sequences (with an average GC content higher than 60%) may result in reduced capture efficiency compared to normal sequences (where the average content of A, T, C, and G bases is 25% each), due to the molecular structure of the C and G bases. For key study regions, such as CGI regions (CpG islands), the increased probe quantity may be necessary to obtain sufficient and accurate CGI data.

[0178] B. Amplicon-based sequencing The transformed ssDNA fragment can be amplified. In some embodiments, amplification is performed using primers designed to anneal to a methylated target sequence having at least one methylation site. Methylation sequencing transformation results in the conversion of unmethylated cytosine to uracil, while 5-methylcytosine remains unaffected. "Target sequence" can refer to a sequence where cytosine known to have a methylation site is fixed as "C" (cytosine), and unmethylated cytosine known to have an unmethylated site is fixed as "U" (uracil); for primer design purposes, "U" can be considered as "T" (thymine).

[0179] In various embodiments, the DNA source is cell-free DNA obtained from whole blood, plasma, serum, or genomic DNA extracted from cells or tissues. In some embodiments, the amplified fragment is between about 100 and 200 base pairs in length. In some embodiments, the DNA source is extracted from a cellular source (e.g., tissue, biopsy, cell line), and the amplified fragment is between about 100 and 350 base pairs in length. In some embodiments, the amplified fragment contains at least one 20-base-pair sequence comprising at least one, at least two, at least three, or more than three CpG dinucleotides. Amplification can be performed using a primer oligonucleotide set according to this disclosure and can be performed using a thermostable polymerase. Amplification of several DNA segments can be performed simultaneously in the same reaction vessel. In some embodiments, two or more fragments are amplified simultaneously. For example, amplification can be performed using polymerase chain reaction (PCR).

[0180] Primers designed to target such sequences can exhibit a bias towards the transformed methylated sequence. In some implementations, PCR primers are designed to be methylation-specific for targeted methylation sequencing applications. Methylation-specific primers can allow for higher sensitivity in some applications. For example, primers can be designed to contain a resolving nucleotide (specific to the methylated sequence after bisulfite conversion) positioned to achieve optimal resolution, for example, in PCR applications. This resolving nucleotide can be located at the last or penultimate 3' position.

[0181] In some implementations, primers are designed to amplify DNA fragments of 75 to 350 base pairs (bp) in length, which is the typical size range for circulating DNA. Optimizing primer design to take target size into account can improve the sensitivity of the methods described herein. Primers can be designed to amplify regions of approximately 50 to 200, approximately 75 to 150, or approximately 100 or 125 bp in length.

[0182] In one implementation, the amplification operation includes using primers containing a unique double index (UDI) sequence.

[0183] In one embodiment, the length of the UDI sequence is greater than or equal to 4 base pairs (bp), 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 11 bp, or 12 bp. In another embodiment, the length of the UDI sequence is less than or equal to 12 bp, 11 bp, 10 bp, 9 bp, 8 bp, 7 bp, 6 bp, 5 bp, or 4 bp.

[0184] In some embodiments, the methylation status of a preselected CpG position within the nucleic acid sequence is detected using an amplicon-based method with methylation-specific PCR (MSP) primer oligonucleotides. The use of methylation-specific primers for amplifying transformed methylated ssDNA allows differentiation between methylated and unmethylated nucleic acids. MSP primer pairs contain at least one primer that hybridizes to the transformed CpG dinucleotide. Therefore, the primer sequence includes at least one CpG, TpG, or CpA dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the 3' position of the C position in the CpG. Therefore, the base sequence of these primers may include a sequence of at least 18 nucleotides in length that hybridizes to the pretreated nucleic acid sequence and its complementary sequence, and the base sequence contains at least one CpG, TpG, or CpA dinucleotide. In some embodiments of the method, the MSP primers have 2 to 5 CpG, TpG, or CpA dinucleotides. In some embodiments, the dinucleotide is located within the 3' half of the primer; for example, for a primer with an 18-base length, the designated dinucleotide is located within the first 9 bases from the 3' end of the sequence. In addition to CpG, TpG, or CpA dinucleotides, the primer may further include several methyl-converted bases (e.g., cytosine to thymine, or guanine to adenosine on the hybrid strand). In some embodiments, the primer is designed to have no more than two cytosine and / or guanine bases.

[0185] In some implementations, each region is amplified using multiple primer pairs. In some implementations, these segments are non-overlapping. Segments can be adjacent or spaced apart (e.g., spaced up to 10 base pairs (bp), 20 bp, 30 bp, 40 bp, or 50 bp apart). Since target regions (including CpG islands, CpG shores, and / or CpG skeletons) are typically longer than 75 to 150 bp, this example allows for the assessment of the methylation status of sites in more portions (or all) of a given target region.

[0186] Primers can be designed to target regions using suitable tools such as Primer3, Primer3Plus, and Primer-BLAST. As discussed, methylation enrichment transformations result in the conversion of unmethylated cytosine to uracil and methylated cytosine to thymine. Therefore, depending on the desired level of methylation specificity, primer localization or targeting can utilize the transformed sequence.

[0187] C. Enzymatic transformation for DNA methylation sequencing applications Bisulfite conversion can impair input DNA and may result in overall yield loss, fragmentation, and biased sequencing data. As an alternative, enzymatic methylation can be used in methylation sequencing workflows. Examples of enzymatic methylation workflows include enzymatic methyl-seq (EM-seq) and TET-assisted pyridineborane sequencing (TAPS).

[0188] EM-seq is a minimally destructive transformational methylation sequencing method for converting cytosine to uracil in nucleic acids. This bisulfite-free method preserves the length of nucleic acid molecules while achieving conversion rates similar to bisulfite sequencing. Furthermore, EM-Seq can result in higher sequencing quality scores for cytosine and guanine base pairs and can provide more uniform coverage of various genomic features, such as CpG islands. EM-Seq comprises two sets of enzymatic reactions. In the initial reaction, deca-11 translocation (TET) enzymes (e.g., TET1, TET2, TET3, Naegleria TET and their engineered forms and / or variants) and β-glucosyltransferases (e.g., T4 BGT) convert 5mC and 5hmC into products that cannot be deaminated by cytosine deaminases (e.g., APOBEC) or deaminase-resistant products. In the second reaction, cytosine deaminases (e.g., APOBEC) deaminize unmodified (e.g., unmethylated) cytosine by converting it to uracil.

[0189] In another implementation, TAPS can be used in an enzymatic methylation sequencing workflow. TAPS is a minimally destructive conversion methylation sequencing method for converting cytosine in nucleic acids to uracil. This bisulfite-free method allows for minimal DNA degradation, thus preserving the length of nucleic acid molecules while achieving conversion rates similar to sodium bisulfite sequencing. TAPS can result in higher sequencing quality scores for cytosine and guanine base pairs and can provide more uniform coverage of various genomic features, such as CpG islands.

[0190] In TAPS, a deca-11 transloase (e.g., TET1) is used to oxidize 5mC and 5hmC to 5caC. Pyridineborane is used to reduce 5caC to dihydrouracil, a uracil derivative, which is subsequently converted to thymine after PCR. TAPS can be performed in two other ways: TAPSβ and chemically assisted pyridineborane sequencing (CAPS). In TAPSβ, β-glucosyltransferase is used to label 5hmC with glucose to protect it from oxidation and reduction reactions and allow for specific detection of 5mC. In CAPS, potassium perruthenate acts as a chemical substitute for Tet1 and specifically oxidizes 5hmC, thus allowing for direct detection.

[0191] In one example of enzymatic methylation, the combination of unmodified C to U enzymatic conversion and an interleaved UMI adaptor consistent with the library insert can be used for targeted sequencing of methylated libraries. For low-depth sequencing applications, this combination allows for a reduction in plasma input volume or cfDNA input compared to bisulfite conversion sequencing because the sample cfDNA is not degraded to the same degree.

[0192] For high-depth sequencing applications, because cfDNA is not degraded to the same degree, higher-depth sequencing can be obtained from similar inputs such as plasma or cfDNA compared to bisulfite conversion sequencing.

[0193] In one instance, the cytosine present in the adaptor nucleic acid is modified with 5-methyl or 5-hydroxymethyl to prevent C-to-T conversion in the adaptor.

[0194] One advantage of this method is that, compared to methods involving bisulfite conversion followed by ssDNA adaptor ligation, pre-conversion adaptor ligation preserves fragment endpoint and length information. Extensive nucleic acid degradation prior to adaptor ligation can lead to the loss of informative fragment endpoint and length information.

[0195] Compared to bisulfite conversion methods, enzymatic conversion of unmodified C to U is less demanding on sample nucleic acid fragments and can result in more complete and uniform coverage. Bisulfite degradation of DNA is heterogeneous, so some sequences are preferentially degraded compared to others, including CG dinucleotides, which are precisely the sites interrogated in methylation sequencing. Therefore, enzymatic methods provide higher CpG site coverage and greater uniformity of captured reads in target enrichment applications compared to bisulfite conversion methods that use the same number of unique reads. Furthermore, non-bisulfite methods (e.g., enzymatic and TAPS-like chemical conversions) offer increased biological signal resolution, particularly the ability to distinguish 5mC and 5hmC methylation in nucleic acid sequences. This information and additional resolution can be valuable in computational and other methods.

[0196] In some instances, subjecting DNA or barcoded DNA to an enzymatic reaction that converts the cytosine nucleobases of DNA or barcoded DNA to uracil nucleobases involves “enzymatic conversion”.

[0197] In various instances, glucosylation and oxidation reactions overcome the observed intrinsic deamination of 5hmC and 5mC by deaminases. Deaminases convert 5mC and unmodified C to U, but not 5ghmC and 5caC. Non-restrictive examples of deaminases include APOBEC (catalytic peptide-like apolipoprotein B mRNA editing enzyme). The embodiments described herein utilize enzymes that are substantially free of sequence bias in the glucosylation, oxidation, and deamination of cytosine. Furthermore, these embodiments provide substantially no non-specific damage to DNA during the glucosylation, oxidation, and deamination reactions.

[0198] In some implementations, a glucosyltransferase (GT) (e.g., β-glucosyltransferase (βGT)) is used to covalently link glucose to 5hmC to protect the modified base from deamination. Other enzymatic or chemical reactions can be used to modify 5hmC to achieve the same effect.

[0199] Typically, and in one aspect, the methods provided herein comprise (a) treating an equal fraction (a portion) of a nucleic acid sample in a reaction mixture with a dioxygenase (e.g., TET2 and βGT) to produce a reaction product in which substantially all modified cytosine (C) is oxidized, or glucosylated in the case of 5hmC; and (b) treating the reaction product with a cytidine deaminase to convert substantially all unmodified C to U. The term “modified” cytosine, as used throughout these examples and embodiments, refers to one or more of 5mC, 5hmC, 5ghmC, 5fC, and 5caC, wherein complete oxidation of 5mC, 5hmC, and 5fC yields 5caC. βGT reacts only with 5hmC. However, some 5hmC may be converted to 5fC by a dioxygenase before glucosylation occurs, and then to 5caC. In the presence of a dioxygenase, most of the 5mC is completely oxidized to 5caC, but some residual 5hmC may be produced. However, residual 5hmC can be glucosylated by βGT to prevent the low deamination rate of 5hmC, which could otherwise reduce the accuracy of methylation sequencing.

[0200] Therefore, the described method largely distinguishes between unmodified and modified cytosine by treating nucleic acids with dioxygenases prior to deamination. However, the amount of naturally occurring 5mC in genomic DNA can significantly exceed the amount of 5hmC, which in turn can exceed the amounts of naturally occurring 5fC and 5caC. Therefore, the amount of naturally occurring modified cytosine is often considered an approximation of the amount of naturally occurring 5mC.

[0201] In one example, the method can be adapted for 5hmC sequencing. The 5hmC sequencing method may further include: treating an equal fraction of a nucleic acid sample with βGT in the absence of dioxygenase, followed by treatment with cytidine deaminase to produce a reaction product in which substantially all 5hmC in the equal fraction is glucosylated, while substantially all unmodified C and 5mC are converted to U. After PCR amplification, U is converted to T, thus making cytosine and 5mC indistinguishable at sequencing. The resulting reaction product can be sequenced and compared with a reference sequence to distinguish 5hmC from C and 5mC. Distinguishing these portions allows mapping these modified nucleotides to a reference sequence, such as a reference sequence from a database or an independently determined reference sequence.

[0202] In some embodiments, the product of the dioxygenase-βGT+deaminase reaction or its amplification product can be sequenced to determine which Cs are methylated (which may include a small subset of 5hmCs) and which Cs are unmodified. In some embodiments, the product of the dioxygenase-free βGT+deaminase reaction or its amplification product can be sequenced to determine which Cs are hydroxymethylated and which Cs are unhydroxymethylated. In some embodiments, the product of the dioxygenase-free βGT+deaminase reaction or its amplification product can be sequenced to determine which Cs are hydroxymethylated and which Cs are unmodified. A reference DNA can be generated by sequencing the resulting reaction product from a nucleic acid sample that has not been reacted with any of the dioxygenase, βGT, and deaminase. Alternatively, the reference sequence is a known reference sequence, such as a reference sequence from a sequence database.

[0203] In one embodiment, the sequence of the product of the dioxygenase reaction with βGT plus deaminase can be compared with a reference sequence. Alternatively, this can also be compared with the sequence of the product of the βGT (without dioxygenase) plus deaminase reaction to determine which cytosines in the nucleic acid sample are methylated relative to hydroxymethyl groups.

[0204] In one aspect, a method is provided for targeted methylation sequencing of cell-free DNA (cfDNA) samples from an object, comprising: a) Link a transformation-tolerant nucleic acid adaptor to a single-stranded nucleic acid molecule in a cfDNA sample, wherein the single-stranded nucleic acid molecule includes untransformed nucleic acids; b) Enzymatically convert unmethylated cytosine in single-stranded nucleic acid molecules into uracil to produce transformed nucleic acids; c) Amplification of transformed nucleic acids via polymerase chain reaction; d) Probe the transformed nucleic acid with a nucleic acid probe complementary to the pre-identified CpG or CH locus to enrich the sequence corresponding to the pre-identified CpG or CH locus. e) Determine the nucleic acid sequence of the transformed nucleic acid at a depth >100x; and f) Compare the nucleic acid sequence of the transformed nucleic acid with the reference nucleic acid sequence of the pre-identified CpG or CH locus to determine the methylation profile of the cfDNA sample from the subject.

[0205] In some embodiments, the transformation-tolerant nucleic acid adaptor is a transformation-resistant nucleic acid adaptor containing one or more deamination-resistant modified cytosines, including but not limited to propynyl-C and pyrrolo-C. In one embodiment, the transformation-resistant adaptor may contain one or more propynyl-C residues. In another embodiment, the transformation-resistant adaptor may contain one or more pyrrolo-C residues. In yet another embodiment, the transformation-resistant adaptor may contain a combination of propynyl-C and pyrrolo-C residues. If the nucleic acid sequence being tested for transformation is a T corresponding to a reference C at the designated CpG locus, then that C is not methylated in the original test nucleic acid fragment. In contrast, if both the nucleic acid sequence being tested for transformation and the reference sequence are C at the designated CpG locus, then that C is methylated in the original test nucleic acid fragment.

[0206] In one instance, the nucleic acid sequence of the transformed nucleic acid molecule is sequenced at a depth of approximately 50-500x, approximately 25-1000x, approximately 50-500x, approximately 250-750x, approximately 500-200x, approximately 750-1500x, or approximately 100-2000x. In some implementations, the nucleic acid sequence is sequenced at a depth >100x or >500x.

[0207] In one instance, the nucleic acid sequence of the transformed nucleic acid molecule was sequenced at a depth of approximately 500x, approximately 1000x, approximately 2000x, approximately 3000x, approximately 4000x, approximately 5000x, approximately 6000x, approximately 7000x, approximately 8000x, approximately 9000x, approximately 10000x, or greater than 5000x.

[0208] In one instance, the nucleic acid sequence of the transformed nucleic acid molecule was sequenced at a depth of approximately 300x uniqueness, approximately 400x uniqueness, approximately 500x uniqueness, approximately 600x uniqueness, approximately 700x uniqueness, approximately 800x uniqueness, approximately 900x uniqueness, or approximately 1000x uniqueness, or greater than 500x uniqueness.

[0209] D. Target enrichment sequencing applications Furthermore, a method for enriching methylated regions of interest in target capture applications during sequencing is provided. A potential problem with using target enrichment capture sets with DNA methylation libraries is a low in-target read ratio / high off-target DNA fragment capture rate. For each region in the set, probes can be programmed to target DNA derived from methylated CpGs or DNA derived from unmethylated CpGs. In either probe type, each CpG site along the region is considered either unmethylated or methylated, as appropriate for the probe type. Probes can hybridize with library molecules after bisulfite / enzymatic conversion and PCR amplification. Only library molecules captured by the probes are subsequently sequenced. This method has the advantage of reducing sequencing costs because only a small fraction of the genome is sequenced. In one instance, approximately 0.1% of the genome was sequenced. In another instance, approximately 0.3% of the genome was sequenced. In another instance, approximately 0.5% of the genome was sequenced. In another instance, approximately 0.7% of the genome was sequenced. In another instance, approximately 1% of the genome was sequenced. In other instances, approximately 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the genome was sequenced. In other instances, approximately 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, or 2% of the genome was sequenced.

[0210] Significant off-target capture rates can occur in both bisulfite- and enzymatically transformed libraries for target capture enrichment. This off-target capture rate is partly due to the C-to-T conversion of all cytosines not present at CpG sites in both types of probes that hybridize with DNA derived from methylated CpG. Reduced cytosine content in the probe leads to decreased sequence complexity, and therefore, lower specificity for probe hybridization with the target library.

[0211] As used herein, the terms "methylated probe" and "unmethylated probe" refer to probes used to hybridize with methylated and unmethylated CpG in transformed nucleic acid sequences, respectively. Probes can be designed to recognize transformed nucleic acid sequences. In transformed methylated CpG probes, C remains C after transformation. In transformed unmethylated CpG probes, C is converted to T after transformation. In both transformed methylated and unmethylated probes, all C values ​​in non-CpG dinucleotides are converted to T after transformation.

[0212] Methylated probes retain some cytosine (e.g., cytosine at the CpG site). In contrast, in unmethylated probes, all cytosine is converted to thymine. Unmethylated probes are less complex than methylated probes and may preferentially contribute to off-target capture rates. In one example, a probe hybridizing to DNA derived from methylated CpG is used in a target enrichment method. In another example, a probe having a sequence substantially complementary to the target that hybridizes to DNA derived from methylated CpG is used in a target enrichment method.

[0213] Probes used for target enrichment that hybridize with DNA derived from methylated CpG can be selected to achieve different aspects. The target capture hybridization reaction occurs at a single temperature. However, the optimal melting temperature (Tm) of probes hybridizing with DNA derived from methylated CpG is, on average, higher than the Tm of probes not designed to hybridize with DNA derived from methylated CpG.

[0214] Cytosine base pairing involves three hydrogen bonds, while thymine base pairing involves two. The conversion of cytosine to thymine in a probe reduces the probe's Tm due to decreased hydrogen bonding. Probes designed to collect DNA derived from methylated fragments can contain more cytosine than probe sets designed to collect unmethylated fragments corresponding to the same genomic region. Therefore, methylated fragment-targeting probes can have a higher Tm relative to their corresponding unmethylated probes. The difference in melting temperature between methylated and unmethylated probes increases with the number of CpG sites in the region. Probes with higher melting temperatures hybridize to target DNA fragments more efficiently than those with lower melting temperatures. Hybridization temperatures are typically chosen to be relatively high to promote mid-target capture. However, at many hybridization temperatures, methylated probes can hybridize more efficiently than unmethylated probes due to the higher melting temperature resulting from the retention of some cytosine. Higher melting temperatures can lead to a bias toward higher CpG methylation levels measured by target capture hybridization methods compared to levels measured by sequencing of pre-captured libraries.

[0215] In one instance, a single probe type, either methylated or unmethylated, was used in the hybridization reaction to enrich highly methylated or hypomethylated library molecules, respectively. Using a single type of methylated or unmethylated probe avoids the problem of different melting temperatures between probe types. Using a single probe type also facilitates more efficient capture (or enrichment) of the same DNA fragment type. In one instance, using a methylated-only probe provided preferential binding to highly methylated regions of interest (ROIs) compared to hypomethylated ROIs. In another instance, using an unmethylated-only probe provided enrichment of unmethylated ROIs.

[0216] Using only a single probe type also allows for higher hybridization temperatures to reduce off-target capture without affecting the relative balance of methylated and unmethylated ROI capture. Therefore, probe sets can be designed according to the desired enrichment of either highly methylated or hypomethylated DNA fragments. In one instance, when quantification of both highly methylated and hypomethylated DNA fragments is required, two parallel but independent hybridization reactions are employed for both methylation states.

[0217] E. Methylation analysis In various instances, enzymatic methylation sequencing results are used to analyze the methylation status of nucleic acids in biological samples. In one instance, whole-genome enzymatic methylation sequencing (“WG EM-seq”) provides high-resolution sequencing by characterizing DNA methylation at virtually every cytidine nucleotide in the genome. Other targeted methods, such as targeted enzymatic methylation sequencing (“TEM-seq”), can be useful for methylation analysis.

[0218] In other instances, assays applicable to bisulfite conversion can be used with minimally invasive conversion methods, such as enzymatic conversion, TAPS, and CAPS. In various instances, assays for methylation analysis can include mass spectrometry, methylation-specific PCR (MSP), reduced representative bisulfite sequencing (RRBS), HELP assays, GLAD-PCR assays, ChIP-on-chip assays, restriction marker genome scanning, methylated DNA immunoprecipitation (MeDIP), pyrosequencing of bisulfite-treated DNA, molecular break light assays, methyl-sensitive Southern blotting, high-resolution melting analysis (HRM or HRMA), ancient DNA methylation reconstruction, or methylation-sensitive single nucleotide primer extension assays (msSNuPE).

[0219] The methylation profile of cfDNA can be identified by applying sequence alignment methods to map methyl-seq reads from the whole genome or targeted methyl sequencing of a human reference genome. Non-limiting examples of sequence alignment methods include bwa-meth, bismark, Last, GSNAP, BSMAP, NovoAlign, Bison, metagenomic phylogenetic analyses (e.g., MetaPhlAn2), BLAT, Burrows-Wheeler Aligner (BWA), Bowtie, Bowtie2, Bfast, BioScope, CLC bio, Cloudburst, Eland / Eland2, GenomeMapper, GnuMap, Karma, MAQ, MOM, Mosaik, MrFAST / MrsFAST, PASS, PerM, RazerS, RMAP, SSAHA2, Segemehl, SeqMap, SHRiMP, Slider / SliderII, Srprism, Stampy, vmatch, ZOOM, and SOAP / SOAP2 alignment tools.

[0220] F. Identification of somatic cell variants In various instances, enzymatically transformed DNA is used to infer the methylation status of C residues in the genome. However, because enzymatic transformation of DNA converts unmethylated C residues to U residues and does not introduce other chemical changes into the DNA, somatic variants that do not correspond to C or T bases in a reference or query sequence can also be identified in the transformed DNA. These somatic variants can be identified using existing methods designed for untransformed DNA.

[0221] G. Inferring nucleosome localization Compared to flanking DNA, cytosine methylation at CpG sites can be significantly enriched across DNA spanning the nucleosome. Therefore, CpG methylation patterns can also be used to infer nucleosome localization using machine learning methods. EM-seq datasets can also be analyzed using the same methods employed for WGS to generate features for machine learning methods and models, regardless of methylation transformation. Subsequently, 5mC patterns can be used to predict nucleosome localization, which can then aid in inferring gene expression and / or disease and cancer classification. In another instance, features can be obtained from a combination of methylation status and nucleosome localization information.

[0222] Metrics used for methylation analysis include, but are not limited to, M-bias (base-wise methylation percentage of CpG, CHG, and CHH), conversion efficiency (e.g., 100-average methylation percentage of CHH), hypomethylated blocks, methylation levels (e.g., global average methylation of CPG, CHH, CHG, chrM, LINE1, or ALU), dinucleotide coverage (normalized dinucleotide coverage), coverage uniformity (e.g., unique CpG sites at 1x and 10x average genome coverage (for S4 runs)), global average CpG coverage (depth), and average coverage at CpG islands, CGI racks, and CGI shores. In one instance, fragment endpoints and length information are used as feature inputs for the analysis. These metrics can also be used as feature inputs for machine learning methods and models.

[0223] In another aspect, this disclosure provides a method comprising: (a) providing a biological sample containing cfDNA from an object; (b) subjecting the cfDNA to conditions sufficient to optionally enrich methylated cfDNA in the sample; (c) enzymatically converting unmethylated cytosine nucleobases of the cfDNA to uracil nucleobases; (d) sequencing the cfDNA to generate sequence reads; (e) computer processing the sequence reads to (i) determine the degree of methylation of the cfDNA based on the presence of uracil nucleobases; and (ii) modeling at least a partial degradation of the cfDNA to generate degradation parameters; and (f) using the degradation parameters and the degree of methylation to determine gene sequence characteristics.

[0224] In some instances, sequencing of cfDNA involves determining the degree of DNA methylation based on the ratio of unconverted to converted cytosine nucleobases. In some instances, converted cytosine nucleobases are detected as uracil nucleobases. In some instances, uracil nucleobases are observed as thymine nucleobases in sequence reads.

[0225] In some instances, generating degradation parameters involves using Bayesian models. In some instances, Bayesian models are based on chain bias or enzymatic conversion or overconversion. In some instances, the computational processing of sequence reads involves using degradation parameters within the framework of pairwise HMMs or Naive Bayes models.

[0226] H. Analysis of differentially methylated regions (DMR) In one instance, methylation analysis is differentially methylated region (DMR) analysis. DMR is used to quantify CpG methylation within genomic regions. These regions are dynamically assigned through discovery. Multiple samples from different categories can be analyzed, and regions with the highest degree of differential methylation between different classifications can be identified. A subset of regions can be selected as differentially methylated and used for classification. The number of CpGs captured in a region can be used for analysis. The size of the region can be variable. In one instance, a pre-discovery process is performed, which bundles multiple CpG sites together as regions. In another instance, DMR is used as input features for processing by machine learning methods and models.

[0227] I. Methylation Haplotype Block and Methylation Haplotype Load In one instance, haplotype block determination is applied to the sample. Identification of methylated haplotype blocks aids in the unconvolution of heterogeneous tissue samples and the mapping of tumor-originating tissues from plasma DNA. Tightly coupled CpG sites, termed methylated haplotype blocks (MHBs), can be identified in WGBS data. An index called methylated haplotype load (MHL) is used for tissue-specific methylation analysis at the block level. This method provides informative blocks that can be used for unconvolution of heterogeneous samples. This method can be used for quantitative estimation of tumor load and mapping of originating tissues in circulating cfDNA. In one instance, haplotype blocks are used as input features for processing by machine learning methods and models.

[0228] J. Targeted methylation interpretation analysis for identifying cell types of origin On one hand, the method is used for targeted methylation interpretation to identify the cell type of origin of cfDNA molecules based on methylation patterns. This method provides a probabilistic model of the co-methylation status of multiple neighboring CpG sites on individual sequencing reads to leverage the universality of DNA methylation for signal amplification. The model develops probabilities for sequencing reads for each cell type, and then a mixed model is developed and fitted to the global cell type distribution.

[0229] Traditional DNA methylation analysis focuses on the methylation rate (β value) of individual CpG sites within a cell population to indicate the proportion of cells with CpG site methylation. This population-average measure is often insufficiently sensitive to capture aberrant methylation signals affecting a small proportion of cfDNA. However, based on the universality of DNA methylation, disease-specific cfDNA reads can be computationally distinguished from normal cfDNA reads.

[0230] Furthermore, given the universality of DNA methylation, the co-methylation status of multiple neighboring CpG sites can be used to easily distinguish cancer-specific cfDNA reads from normal cfDNA reads. The average methylation value of all CpG sites in a given read (denoted as α value) provides the difference (0 and 1) between aberrantly methylated cfDNA and normal cfDNA (α tumor = 0% and α normal = 100%). The methylation α value is used to estimate whether the co-methylation probability of all CpG sites in a read follows the DNA methylation signature of a disease. This method can sensitively identify cfDNA of multiple cell types of origin from all cfDNA in plasma.

[0231] In various instances, alignment tools are used to align reads to a reference genome and interpret methylated cytosine. PCR duplicates are removed, and the number of methylated and unmethylated cytosines at each CpG site is quantified. The methylation level of a CpG cluster is calculated as the ratio between the number of methylated cytosines and the total number of cytosines within the cluster. This WGBS data processing procedure calculates the mean methylation level of CpG clusters in normal plasma samples used to identify methylation biomarkers. When plasma cfDNA samples are used as test data, the co-methylation status of all CpG sites in individual sequencing reads aligned to the biomarker group is extracted and then processed by a machine learning model. In this approach, CpG methylation interpretation is used as input features for methylation status analysis and feature generation. To improve input data quality for high-coverage cfDNA methylation data, reads covering <2, <3, or <4 CpG sites can be filtered out.

[0232] The methylation sequencing methods described in this paper improve sequencing read quality, for example, by reducing PCR errors and bias and minimizing DNA degradation that occurs during bisulfite conversion. In one instance, methylation sequencing data is used to model overlapping regions. In another instance, machine learning modeling can determine the cell type of origin for identified methylated DNA regions.

[0233] In various instances, the model can classify more than two cell types of origin. In other instances, the model can classify sequences into 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 50, 75, 100, or more than 100 different cell types.

[0234] K. DNA hydroxymethylation analysis In one aspect of the present invention, 5hmC sequencing can be implemented as follows: hydroxymethylation in the adaptor nucleic acid is substituted during adaptor ligation, and glucose is conjugated to 5hmC residues in the insert fragment of the test nucleic acid library using only βGT, without using dioxygenase and βGT to conjugate 5mC and 5hmC. When the resulting sequencing data is compared with a reference genome, each C position corresponding to a C in the test sequence shown in the reference is interpreted as a hydroxymethylated C, and each C position corresponding to a T in the test sequence shown in the reference is interpreted as an unmodified C or a methylated C. Therefore, the data interpretation for hydroxymethylation analysis is the same as that for methylation analysis.

[0235] In one aspect of the present invention, methylated and hydroxymethylated sequencing libraries can be compared to indicate the level of each cytosine modification (e.g., 5m or 5mC) at single nucleotide resolution.

[0236] In one aspect of the present invention, since the hydroxymethylation status readout is the same as the methylation status, all analytical methods used for methylation sequencing data can be applied to hydroxymethylation sequencing data.

[0237] III. Computer Systems and Machine Learning Methods A. Sample characteristics As used in this article, in the context of machine learning and pattern recognition, the term "feature" can refer to a single, measurable characteristic or feature of an observed phenomenon. Features are typically numerical, but structural features (such as strings and graphs) are used for syntactic pattern recognition. The concept of "feature" is related to the concept of explanatory variables used in statistical techniques such as linear regression.

[0238] In one implementation, the features are processed into a feature matrix for machine learning analysis.

[0239] For multiple measurements, the system identifies a feature set for processing using a machine learning model. The system performs measurements on each molecular category and generates a feature vector based on the measurements. The system then processes the feature vectors using a machine learning model to obtain an output classification indicating whether the biological sample possesses a specified characteristic.

[0240] In one implementation, the machine learning model outputs a classifier that distinguishes individuals or features from two groups or categories within a group or set of features. In one implementation, the classifier is a trained machine learning classifier.

[0241] In one implementation, informative loci or traits of biomarkers in cancer tissue are determined to form a profile. Receiver operating characteristic (ROC) curves can be used to plot the performance of a specific trait (e.g., any of the biomarkers described herein and / or any additional biomedical information items) in distinguishing two populations (e.g., individuals who respond to and do not respond to a treatment agent). In some implementations, trait data for the entire population (e.g., cases and controls) are sorted in ascending order based on the values ​​of individual traits.

[0242] In some implementation schemes, the condition is advanced adenoma (AA), colorectal cancer (CRC), colorectal epithelial cancer, or inflammatory bowel disease.

[0243] The term "input feature" or "feature" refers to variables used by a model to predict the output classification (label) of a sample, such as conditions, sequence content (e.g., mutations), suggested data collection practices, or suggested treatments. Variable values ​​can be determined for a sample and used to determine the classification. Examples of input features for genetic data include alignment variables related to the alignment of sequence data (e.g., sequence reads) to the genome, and non-alignment variables, such as those related to the sequence content of the sequence reads, measurements of proteins or autoantibodies, or the average methylation level of a genomic region.

[0244] Variable values ​​for samples can be determined and used to determine classification. Examples of input features for genetic data include alignment variables related to the alignment of sequence data (e.g., sequence reads) with the genome, and non-alignment variables, such as those related to the sequence content of sequence reads, measurements of proteins or autoantibodies, or the average methylation level of genomic regions. In various instances, genetic features such as V-map measurements, FREE-C, cfDNA measurements at transcription start sites, and DNA methylation levels on cfDNA fragments are used as input features for machine learning methods and models.

[0245] In one instance, sequencing information includes information about multiple genetic characteristics, such as, but not limited to, transcription start sites, transcription factor binding sites, chromatin open and closed states, nucleosome localization or occupancy, etc.

[0246] B. Data Analysis In some embodiments, this disclosure provides systems, methods, or kits for implementing data analysis in software applications, computing hardware, or both. In various embodiments, the analytical application or system includes at least a data receiving module, a data preprocessing module, a data analysis module (which can operate on one or more types of genomic data), a data interpretation module, or a data visualization module. In one embodiment, the data receiving module may include a computer system that connects laboratory hardware or instruments to a computer system that processes laboratory data. In one embodiment, the data preprocessing module may include a hardware system or computer software that operates on the data to prepare it for analysis. Examples of operations that can be applied to the data in the preprocessing module include affine transformations, denoising operations, data cleaning, reformatting, or subsampling. The data analysis module may be specifically designed to analyze genomic data from one or more genomic materials, which may, for example, acquire assembled genomic sequences and perform probabilistic and statistical analyses to identify anomalous patterns associated with disease, pathology, state, risk, condition, or phenotype. The data interpretation module may use analytical methods, for example, from statistics, mathematics, or biology, to support the understanding of the relationship between the identified anomalous patterns and health status, functional state, prognosis, or risk. The data visualization module can use mathematical modeling, computer graphics, or rendering methods to create visual representations of data, which can facilitate the understanding or interpretation of results.

[0247] In various implementations, machine learning methods are applied to distinguish samples within a sample population. In one implementation, machine learning methods are applied to distinguish between healthy samples and late-stage adenoma samples.

[0248] In one implementation, one or more machine learning operations used to train the methylation-based prediction engine include one or more of the following: generalized linear models, generalized additive models, nonparametric regression operations, random forest classifiers, spatial regression operations, Bayesian regression models, time series analysis, Bayesian networks, Gaussian networks, decision tree learning operations, artificial neural networks, recurrent neural networks, reinforcement learning operations, linear / nonlinear regression operations, support vector machines, clustering operations, and genetic algorithm operations.

[0249] In various implementation schemes, computer processing methods are selected from logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed random neighborhood embedding (t-SNE), multilayer perceptron (MLP), network clustering, neurofuzzy logic, and artificial neural networks.

[0250] In some embodiments, the methods disclosed herein may include computational analysis of nucleic acid sequencing data from samples from an individual or from multiple individuals. The analysis may identify variants inferred from the sequence data to identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analytical methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis functions, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline variation or somatic mutation. In some embodiments, a variant may refer to a known variant. A known variant may be scientifically validated or reported in the literature. In some embodiments, a variant may refer to a presumptive variant associated with a biological change. The biological change may be known or unknown. In some embodiments, a presumptive variant may be reported in the literature but has not yet been biologically validated.

[0251] Alternatively, presumed variants not reported in the literature can be inferred based on the computational analyses disclosed herein. In some implementations, germline variants may refer to nucleic acids that induce natural or normal variations.

[0252] Natural or normal variations can include, for example, skin color, hair color, and normal weight. In some embodiments, somatic mutations can refer to nucleic acids that induce acquired or aberrant variations. Acquired or aberrant variations can include, for example, cancer, obesity, conditions, symptoms, diseases, and ailments. In some embodiments, the analysis can include distinguishing between germline variants. Germline variants can include, for example, proprietary variants and somatic mutations. In some embodiments, clinicians or other healthcare professionals can use the identified variants to improve healthcare approaches, diagnostic accuracy, and reduce costs.

[0253] This article also provides improved methods and computational systems or software media that can distinguish between sequence errors, somatic mutations, and germline variants in nucleic acids introduced through amplification and / or sequencing technologies. The provided methods can include simultaneous interpretation and scoring of variants from aligned sequencing data of all samples obtained from the patient.

[0254] Samples obtained from subjects other than patients may also be used. Additional samples may also be collected from subjects previously analyzed by sequencing assays or targeted sequencing assays (e.g., targeted resequencing assays). The methods, computational systems, or software media disclosed herein can improve the identification and accuracy of variants or mutations (e.g., germline or somatic, including copy number variations, single nucleotide variations, insertions / deletions, gene fusions) and lower the detection limit by reducing the number of false positives and false negatives identified.

[0255] C. Classifier generation In one aspect, this system and method provide a classifier generated based on feature information derived from methylation sequence analysis of cfDNA biological samples prepared using the ssDNA library preparation method described herein. The classifier forms part of a prediction engine for distinguishing groups within a population based on methylation sequence features identified in the biological sample (e.g., cfDNA).

[0256] In one implementation, a classifier is created by normalizing methylation information by formatting similar portions of the methylation information into a uniform format and scale; the normalized methylation information is stored in a columnar database; a methylation prediction engine is trained by applying one or more machine learning operations to the stored normalized methylation information, the methylation prediction engine mapping combinations of one or more features for a specific population; the methylation prediction engine is applied to accessed field information to identify group-related methylations; and individuals are classified into groups.

[0257] Specificity can be defined as the probability of testing negative in a disease-free population. Specificity equals the number of disease-free individuals who test negative divided by the total number of disease-free individuals.

[0258] In various implementation schemes, the model, classifier, or prediction test has a specificity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 97%, at least 98%, or at least 99%.

[0259] Sensitivity can be defined as the probability of a positive test in a population with a disease. Sensitivity equals the number of infected individuals who test positive divided by the total number of infected individuals.

[0260] In various implementation schemes, the model, classifier, or prediction test has a sensitivity of at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 97%, at least 98%, or at least 99%.

[0261] In one implementation plan, the groups are selected from healthy (asymptomatic), cancer, bowel-related diseases, immune-mediated inflammatory diseases, neurological diseases, kidney diseases, prenatal diseases, and metabolic diseases.

[0262] D. Digital processing equipment In some embodiments, the subject matter described herein may include a digital processing device or its use. In some embodiments, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform device functions. In some embodiments, the digital processing device may include an operating system configured to execute executable instructions. In some embodiments, the digital processing device may be optionally connected to a computer network. In some embodiments, the digital processing device may be optionally connected to the Internet, enabling it to access the World Wide Web. In some embodiments, the digital processing device may be optionally connected to a cloud computing infrastructure. In some embodiments, the digital processing device may be optionally connected to an intranet. In some embodiments, the digital processing device may be optionally connected to a data storage device.

[0263] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netboard computers, set-top computers, handheld computers, internet devices, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, tablet computers with readers, thin tablets, and convertible configurations.

[0264] In some implementations, the digital processing device may include an operating system configured to execute executable instructions. For example, the operating system may include software, including programs and data, that manages the device's hardware and provides services for executing applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, and NetBSD. Linux, Apple Mac OS X Server Oracle Solaris Windows Server and Novell NetWare Non-limiting examples of suitable personal computer operating systems include Microsoft. Windows Apple MacOS X UNIX and UNIX-like operating systems, such as GNU / Linux In some implementations, the operating system may be provided by cloud computing, and cloud computing resources may be provided by one or more service providers.

[0265] In some embodiments, the device may include a storage and / or memory device. A storage and / or memory device may be one or more physical means for temporarily or permanently storing data or programs. In some embodiments, the device may be volatile memory and requires power to maintain the stored information. In some embodiments, the device may be non-volatile memory and retain the stored information when the digital processing device is not powered. In some embodiments, non-volatile memory may include flash memory. In some embodiments, non-volatile memory may include dynamic random access memory (DRAM). In some embodiments, non-volatile memory may include ferroelectric random access memory (FRAM). In some embodiments, non-volatile memory may include phase-change random access memory (PRAM). In some embodiments, the device may be a storage device, including, for example, CD-ROMs, DVDs, flash memory devices, disk drives, tape drives, optical disc drives, and cloud-based storage. In some embodiments, the storage and / or memory device may be a combination of devices such as those disclosed herein.

[0266] In some embodiments, the digital processing device may include a display to send visual information to a user. In some embodiments, the display may be a cathode ray tube (CRT). In some embodiments, the display may be a liquid crystal display (LCD). In some embodiments, the display may be a thin-film transistor liquid crystal display (TFT-LCD). In some embodiments, the display may be an organic light-emitting diode (OLED) display. In some embodiments, the OLED display may be a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some embodiments, the display may be a plasma display. In some embodiments, the display may be a video projector. In some embodiments, the display may be a combination of devices such as those disclosed herein.

[0267] In some embodiments, the digital processing device may include an input device to receive information from a user. In some embodiments, the input device may be a keyboard. In some embodiments, the input device may be a pointing device, including, for example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some embodiments, the input device may be a touchscreen or multi-touchscreen. In some embodiments, the input device may be a microphone to capture voice or other sound input. In some embodiments, the input device may be a camera for capturing motion or visual input. In some embodiments, the input device may be a combination of devices such as those disclosed herein.

[0268] E. Non-transitory computer-readable storage medium In some embodiments, the subject matter disclosed herein may include one or more non-transitory computer-readable storage media encoded with programs, said programs including instructions executable by an operating system of an optionally networked digital processing device. In some embodiments, the computer-readable storage medium may be a tangible component of the digital processing device. In some embodiments, the computer-readable storage medium may be optionally removable from the digital processing device. In some embodiments, the computer-readable storage medium may include, for example, CD-ROMs, DVDs, flash memory devices, solid-state storage, disk drives, tape drives, optical disc drives, cloud computing systems and services, etc. In some embodiments, the programs and instructions may be permanently, substantially permanently, semi-permanently, or non-transitory encoded on the medium.

[0269] F. Computer System This disclosure provides a computer system that is programmed to implement the methods of this disclosure. Figure 11 A computer system 1101 is shown, which is programmed or otherwise configured to store, process, identify, or interpret patient data, biological data, biological sequences, or reference sequences. The computer system 1101 can process various aspects of the patient data, biological data, biological sequences, or reference sequences disclosed herein. The computer system 1101 can be a user's electronic device or a computer system located remotely relative to an electronic device. The electronic device can be a mobile electronic device.

[0270] Computer system 1101 includes a central processing unit (CPU, also referred to herein as a “processor” and “computer processor”) 1105, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 1101 also includes memory or memory location 1110 (e.g., random access memory, read-only memory, flash memory), electronic storage unit 1115 (e.g., hard disk), communication interface 1120 for communicating with one or more other systems (e.g., network adapter), and peripheral devices 1125, such as cache, other memory, data storage, and / or electronic display adapters. Memory 1110, storage unit 1115, interface 1120, and peripheral devices 1125 communicate with CPU 1105 via a communication bus (solid line) such as a motherboard. Storage unit 1115 may be a data storage unit (or data repository) for storing data. Computer system 1101 may be operatively coupled to computer network (“network”) 1130 by means of communication interface 1120. Network 1130 may be the Internet, the Internet, and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some embodiments, network 1130 is a telecommunications and / or data network. Network 1130 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some embodiments, with the assistance of computer system 1101, network 1130 can implement a peer-to-peer network, which allows devices coupled to computer system 1101 to act as clients or servers.

[0271] CPU 1105 can execute machine-readable instruction sequences, which can be embodied in a program or software. The instructions can be stored in a memory location, such as memory 1110. These instructions can be directed to CPU 1105, which can then program or otherwise configure CPU 1105 to implement the methods of this disclosure. Examples of operations performed by CPU 1105 can include fetching, decoding, executing, and writing back.

[0272] CPU 1105 may be part of a circuit, such as part of an integrated circuit. One or more other components of system 1101 may be included in the circuit. In some embodiments, the circuit is an application-specific integrated circuit (ASIC).

[0273] Storage unit 1115 may store files, such as drivers, libraries, and saved programs. Storage unit 1115 may store user data, such as user preferences and user programs. In some embodiments, computer system 1101 may include one or more additional data storage units external to computer system 1101, such as those located on a remote server communicating with computer system 1101 via an intranet or the Internet.

[0274] Computer system 1101 can communicate with one or more remote computer systems via network 1130. For example, computer system 1101 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), thin tablets, or tablet PCs (e.g., Apple). iPad, Samsung Galaxy Tab), telephone, smartphone (e.g., Apple) iPhone, Android-enabled devices, Blackberry (or personal digital assistant). Users can access computer system 1101 via network 1130.

[0275] The methods described herein can be implemented using machine-executable code (e.g., a computer processor) stored in an electronic storage location (e.g., on memory 1110 or electronic storage unit 1115) of computer system 1101. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by processor 1105. In some embodiments, the code can be retrieved from storage unit 1115 and stored on memory 1110 for access by processor 1105 at any time. In some embodiments, electronic storage unit 1115 can be excluded, and machine-executable instructions are stored on memory 1110.

[0276] The code can be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or it can be interpreted or compiled during runtime. The code can be provided in a programming language, which can be selected to enable the code to be executed in a pre-compiled, interpreted, or compiled manner.

[0277] Aspects of the systems and methods provided herein (e.g., computer system 1101) can be embodied in a programmable manner. These aspects of the technology can be considered as “products” or “manufactured articles” typically in the form of machine (or processor) executable code and / or associated data, carried or embodied in a type of machine-readable medium. Machine-executable code can be stored on electronic storage units, such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. “Storage” type media can include any or all tangible memory of a computer, processor, or similar device, or related modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, for example, from a management server or host computer to a computer platform for an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used between physical interfaces between local devices, via wired and fiber optic terrestrial networks, and via various air links. Physical elements carrying such waves (e.g., wired or wireless links, optical links, etc.) can also be considered as media carrying software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as "computer or machine readable medium" refer to any medium involved in providing execution instructions to the processor.

[0278] Therefore, machine-readable media (such as computer-executable code) can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, such as any storage device or similar device in any computer, such as those that can be used to implement the database shown in the figure. Volatile storage media include dynamic memory, such as the main memory of such computer platforms. Tangible transmission media include coaxial cables, copper wires, and optical fibers, including wires that form the bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card tapes, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or cartridges, carrier-transmitted data or instructions, cables or links that transmit such carriers, or any other medium that a computer can read programming code and / or data. Many of these forms of computer-readable media may involve transmitting one or more sequences of one or more instructions to a processor for execution.

[0279] Computer system 1101 may include or communicate with an electronic display 1135, the electronic display 1135 including a user interface (UI) 1140 for providing, for example, nucleic acid sequences, enriched nucleic acid samples, expression profiles, and analysis of expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0280] The methods and systems disclosed herein can be implemented by one or more algorithms. The algorithms can be implemented by software executed by the central processing unit 1105. For example, the algorithms can probe multiple regulatory elements, sequence nucleic acid samples, enrich nucleic acid samples, determine the expression profile of nucleic acid samples, analyze the expression profile of nucleic acid samples, and archive or disseminate the analysis results of the expression profile.

[0281] In some implementations, the subject matter disclosed herein may include at least one computer program or its purpose. A computer program may be a sequence of instructions executable in a digital processing device, such as a CPU, GPU, or TPU, written to perform a specified task. Computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), data structures, etc., that perform a specific task or implement a specific abstract data type. Computer programs can be written in various versions of various languages.

[0282] The functionality of computer-readable instructions can be combined or assigned as needed in various environments. In some embodiments, a computer program may include a sequence of instructions. In some embodiments, a computer program may include multiple sequences of instructions. In some embodiments, a computer program may be provided from one location. In some embodiments, a computer program may be provided from multiple locations. In some embodiments, a computer program may include one or more software modules. In some embodiments, a computer program may include, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ins, or add-ons, or combinations thereof.

[0283] In some implementations, the computer processing can be a statistical, mathematical, biological, or any combination thereof method. In some implementations, the computer processing method includes dimensionality reduction methods, including, for example, logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis functions, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, network clustering, and neural networks.

[0284] In some implementations, the computer processing method is a supervised machine learning method, including, for example, regression, support vector machines, tree-based methods, and networks.

[0285] In some implementations, the computer processing method is an unsupervised machine learning method, including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0286] G. Database In some embodiments, the subject matter disclosed herein may include one or more databases, or their use for storing patient data, biological data, biological sequences, or reference sequences. Reference sequences may be derived from a database. Various databases may be suitable for storing and retrieving sequence information. In some embodiments, suitable databases may include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relational model databases, association databases, and XML databases. In some embodiments, the database may be Internet-based. In some embodiments, the database may be web-based. In some embodiments, the database may be cloud-based. In some embodiments, the database may be based on one or more local computer storage devices.

[0287] IV. Cancer Detection and Diagnosis The trained machine learning methods, models, and discriminant classifiers described in this paper can be used in a variety of medical applications, including cancer detection, diagnosis, and treatment responsiveness. Because the models are trained using individual metadata and analyte-derived features, the applications can be tailored to stratify individuals within a population and guide treatment decisions accordingly.

[0288] A. Detection / Diagnosis The methods and systems presented herein can use artificial intelligence-based approaches to perform predictive analytics to analyze data obtained from an object (patient) to generate detection and / or diagnostic outputs for an object suffering from cancer (e.g., CRC) or other indications. For example, an application can apply predictive algorithms to the acquired data to generate a cancer detection, thereby providing a diagnosis that the object has cancer. The predictive algorithms can include artificial intelligence-based predictors, such as machine learning-based predictors, configured to process the acquired data to generate a diagnosis of cancer in an object.

[0289] A machine learning predictor can be trained by using a dataset from one or more patient cohorts with cancer (e.g., a dataset generated by multianalytical assays of biological samples from individuals) as input and the known diagnostic outcome (e.g., stage and / or tumor score) of the subject as the output of the machine learning predictor.

[0290] Training datasets (e.g., datasets generated by multianalytical assays of biological samples from individuals) can be generated from one or more sets of objects, for example, that share common characteristics (features) and outcomes (labels). Training datasets may include a set of features and diagnostically relevant labels corresponding to those features. Features may include characteristics, such as, for example, certain ranges or categories of cfDNA assay measurements, such as the counts of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap with or fall within each of a reference genomic bin (genomic window) set. For example, a set of features collected from a given object at a given time point can be used collectively as a diagnostic marker that can indicate the identified cancer in the object at that time point. Features may also include labels indicating the diagnostic outcome of the object, such as a diagnostic outcome for one or more cancers.

[0291] The markers may include outcomes, such as, for example, the subject's known diagnostic outcome (e.g., stage and / or tumor score). Outcomes may include cancer-related characteristics of the subject. For example, a characteristic could indicate that the subject has one or more types of cancer.

[0292] The training set (e.g., training dataset) can be selected by randomly sampling datasets corresponding to one or more object sets (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). Alternatively, the training set (e.g., training dataset) can be selected by proportionally sampling datasets corresponding to one or more object sets (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). The training set can be balanced among datasets corresponding to one or more object sets (e.g., patients from different clinical sites or trials). The machine learning predictor can be trained until certain predetermined conditions of accuracy or performance are met, such as having a minimum expected value corresponding to a diagnostic accuracy metric. For example, a diagnostic accuracy metric could correspond to the prediction of the diagnosis, stage, or tumor score of one or more cancers in the object set.

[0293] Examples of measures of detection and diagnostic accuracy may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the ROC curve (AUC) corresponding to diagnostic accuracy in detecting or predicting cancer (e.g., colorectal cancer).

[0294] In another aspect, this disclosure provides a method for detecting or identifying cancer in a subject, comprising: (a) providing a biological sample containing an ssDNA molecule derived from a cell-free DNA sample of the subject; (b) performing methylation sequencing on the ssDNA molecule from the subject to generate a plurality of sequencing reads; (c) aligning the sequencing reads to a reference genome; (d) generating quantitative measures of the sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 distinct regions, each of the at least about 10 distinct regions; and (e) applying a trained algorithm to the first feature set to generate the probability that the subject has the cancer.

[0295] For example, such predetermined conditions could be values ​​for the sensitivity of predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer), including, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0296] As another example, such predetermined conditions could be values ​​that predict the specificity of cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer), including, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0297] As another example, such a predetermined condition could be a positive predictive value (PPV) for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer) including, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0298] As another example, such a predetermined condition could be a negative predictive value (NPV) for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer) including, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0299] As another example, such a predetermined condition could be that the AUC of the ROC curve predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer) includes values ​​of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0300] In some instances of any of the foregoing aspects, the method further includes monitoring the progression of a disease in the subject, wherein the monitoring is based at least in part on genetic sequence characteristics. In some instances, the disease is cancer.

[0301] In some instances of any of the foregoing aspects, the method further includes determining the tissue of origin of the cancer in the object, wherein the determination is based at least in part on genetic sequence characteristics.

[0302] In some instances of any of the foregoing aspects, the method further includes estimating the tumor burden in the subject, wherein the estimation is based at least in part on genetic sequence characteristics.

[0303] B. Treatment responsiveness The predictive classifiers, systems, and methods described herein can be used to classify populations of individuals for a variety of clinical applications (e.g., based on multianalyte assays of an individual's biological sample). Examples of such clinical applications include detecting early-stage cancer, diagnosing cancer, classifying cancer into specific disease stages, or determining responsiveness or resistance to therapeutic agents for treating cancer.

[0304] The methods and systems described herein are applicable to various cancer types, similar to grading and staging, and are therefore not limited to a single cancer disease type. Therefore, combinations of analytes and assays can be used in this system and method to predict cancer treatment responsiveness across different cancer types in different tissues and to classify individuals based on treatment responsiveness. In one example, the classifier described herein stratifies individuals into treatment responders and non-responders.

[0305] This disclosure also provides a method for identifying drug targets (e.g., for a specific category-related / important gene) for a condition or disease of interest, comprising assessing the gene expression level of at least one gene in a sample obtained from an individual; and using a neighborhood analysis method to identify genes that are relevant to the classification of the sample, thereby identifying one or more drug targets that are relevant to the classification.

[0306] This disclosure also provides a method for determining the efficacy of a drug designed to treat a disease category, comprising obtaining a sample from an individual suffering from the disease category; subjecting the sample to the drug; assessing the gene expression level of at least one gene in the drug-exposed sample; and using a computer model built based on a weighted voting scheme to classify the drug-exposed sample into the disease category according to the relative gene expression level of the sample relative to the model.

[0307] This disclosure also provides a method for determining the efficacy of a drug designed to treat a disease category, wherein an individual has been treated with the drug, the method comprising obtaining a sample from the individual treated with the drug; assessing the gene expression level of at least one gene in the sample; and classifying the sample into the disease category using a model built based on a weighted voting scheme, including evaluating the gene expression level of the sample against the gene expression level of the model.

[0308] Another application is a method for determining whether an individual belongs to a phenotypic category (e.g., intelligence, treatment response, lifespan, likelihood of viral infection, or obesity), which includes obtaining a sample from the individual; assessing the gene expression level of at least one gene in the sample; and classifying the sample into the disease category using a model built based on a weighted voting scheme, including evaluating the gene expression level of the sample against the gene expression level of the model.

[0309] Biomarkers can be used to predict the prognosis of colorectal cancer patients. The ability to classify patients as high-risk (poor prognosis) or low-risk (favorable prognosis) enables the selection of appropriate treatments for these patients. For example, high-risk patients may benefit from aggressive therapies, while for low-risk patients, therapies may not offer significant advantages.

[0310] Predictive biomarkers can guide treatment decisions by identifying subsets of patients who may be “superior responders” to a particular cancer therapy or individuals who may benefit from alternative treatments.

[0311] On the one hand, the systems and methods described herein that involve classifying populations based on treatment responsiveness refer to cancers treated with chemotherapeutic agents of the following categories: DNA damaging agents, DNA repair targeted therapies, DNA damage signaling inhibitors, inhibitors of DNA damage-induced cell cycle arrest, and inhibition of processes that indirectly lead to DNA damage, but are not limited to these categories. Each of these chemotherapeutic agents can be considered a “DNA damage therapeutic agent.”

[0312] Patient analyte data are categorized into high-risk and low-risk patient groups, such as patients with high-risk or low-risk clinical recurrence, and the results can be used to determine treatment duration. For example, patients identified as high-risk may receive adjuvant chemotherapy after surgery. Patients considered low-risk may not receive adjuvant chemotherapy after surgery. Therefore, this disclosure provides, in some aspects, a method for preparing gene expression profiles of colorectal cancer tumors indicating recurrence risk.

[0313] In various instances, the classifier described in this paper stratifies groups of individuals into treatment responders and non-responders.

[0314] In various instances, treatments have been selected from alkylating agents, plant alkaloids, antitumor antibiotics, antimetabolites, topoisomerase inhibitors, retinoids, checkpoint inhibitor therapy, and VEGF inhibitors.

[0315] Examples of treatments that can stratify a population into responders and non-responders include, but are not limited to: chemotherapy agents, including sorafenib, regorafenib, imatinib, eribulin, gemcitabine, capecitabine, pazopanib, lapatinib, dabrafenib, sunitinib, crizotinib, everolimus, torisirolimus, sirolimus, axitinib, gefitinib, anastrozole, bicalutamide, fulvestrant, raltitrexed, pemetrexed, and goserelin acetate. Acetate, erlotinib, vemurafenib, vismodegib, tamoxifen citrate, paclitaxel, docetaxel, cabazitaxel, oxaliplatin, ziv-aflibercept, bevacizumab, trastuzumab, pertuzumab, panitumumab, taxane, bleomycin, melphalan, plumbagin, camptosar, mitomycin C, mitoxantrone, poly(styrene-maleic acid)-conjugated neo-oncogenes. neocarzinostatin (SMANCS), doxorubicin, pegylated doxorubicindoxorubicin, FOLFORI, 5-fluorouracil, temozolomide, palireotide, tegafur, gimeracil, oteracil, itraconazole, bortezomib, lenalidomide, irinotecan, epirubicin, romidepsin, resminostat, tasquinimod, refamitinib, lapatinib, Tyverb Arenegyr, NGR-TNF, pasireotide, Signifor ticilimumab, tremelimumab, lansoprazole, PrevOnco ABT-869, Linifanib, Voronanib, Tivantinib, Tarceva Erlotinib, Stivarga Regorafenib, Fluorosorafenib, Brivanib, Liposome Doxorubicin, Lenvatinib, Ramucirumab, Peretinoin, Muparfostat, Teysuno Tegafur, gimeracil, oteracil, and orantinib; and antibody therapies, but not limited to alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, or rituximab.

[0316] In other instances, populations can be stratified into responders and non-responders based on checkpoint inhibitor therapies (such as compounds that bind to PD-1 or CTLA4).

[0317] In other instances, populations can be stratified into responders and non-responders based on anti-VEGF therapies that bind to VEGF pathway targets.

[0318] V. Indications In some instances, a biological condition may include a disease. In some instances, a biological condition may be a stage of a disease. In some instances, a biological condition may be a gradual change in a biological state. In some instances, a biological condition may be a treatment effect. In some instances, a biological condition may be a drug effect. In some instances, a biological condition may be a surgical effect. In some instances, a biological condition may be a biological state resulting from lifestyle changes. Non-limiting examples of lifestyle changes include changes in diet, smoking, and sleep patterns. In some instances, the biological condition is unknown. The analysis described herein may include machine learning to infer or interpret unknown biological conditions.

[0319] In one instance, this system and method are specifically used for applications related to colon cancer: cancer that forms in the tissue of the colon (the longest part of the large intestine). Most colon cancers are adenocarcinomas (cancers that originate from cells that form the lining of internal organs and have gland-like characteristics). Cancer progression is characterized by the stage or extent of cancer in the body. Staging is usually based on the size of the tumor, whether lymph nodes contain cancer, and whether the cancer has spread from the primary site to other parts of the body. Colon cancer is staged as stages I, II, III, and IV. Unless otherwise stated, the term "colon cancer" refers to colon cancer in stages 0, I, II (including IIA or IIB), III (including IIIA, IIIB, or IIIC), or IV. In some instances described herein, the colon cancer is from any stage. In one instance, the colon cancer is stage I colorectal cancer. In one instance, the colon cancer is stage II colorectal cancer. In one instance, the colon cancer is stage III colorectal cancer. In one instance, the colon cancer is stage IV colorectal cancer.

[0320] Conditions that can be inferred using the disclosed methods include, for example, cancer, bowel-related diseases, immune-mediated inflammatory diseases, neurological diseases, kidney diseases, prenatal diseases, and metabolic diseases.

[0321] In some instances, the methods disclosed herein can be used to diagnose cancer. Non-limiting examples of cancer include adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal epithelial carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0322] Non-limiting examples of cancers that can be inferred through the disclosed methods and systems include acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brainstem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medullary epithelioma, pineal parenchymal tumor, breast cancer, bronchial tumor, Burkitt lymphoma, non-Hodgkin lymphoma, carcinoid tumors, cervical cancer, chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, and Ewing sarcoma. Sarcoma, ocular cancer, intraocular melanoma, retinoblastoma, fibrous histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular carcinoma, Hodgkin lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell carcinoma, small cell carcinoma, melanoma, oral cancer, myelodysplastic syndrome, multiple myeloma, medulloblastoma, nasal cavity cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal carcinoma, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papilloma, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumor, plasma cell vegetation, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome Syndrome), skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, laryngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor.

[0323] Non-limiting examples of bowel-related diseases that can be inferred through the disclosed methods and systems include Crohn's disease, colitis, ulcerative colitis (UC), inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), and celiac disease. In some instances, the diseases are inflammatory bowel disease, colitis, ulcerative colitis, Crohn's disease, microscopic colitis, collagenous colitis, lymphocytic colitis, diversion colitis, Behçet's disease, and indeterminate colitis.

[0324] Non-limiting examples of immune-mediated inflammatory diseases that can be inferred through the disclosed methods and systems include psoriasis, sarcoidosis, rheumatoid arthritis, asthma, rhinitis (hay fever), food allergies, eczema, lupus, multiple sclerosis, fibromyalgia, type 1 diabetes, and Lyme disease. Non-limiting examples of neurological diseases that can be inferred through the disclosed methods and systems include Parkinson's disease, Huntington's disease, multiple sclerosis, Alzheimer's disease, stroke, epilepsy, neurodegeneration, and neuropathy.

[0325] Non-limiting examples of kidney diseases that can be inferred using the disclosed methods and systems include interstitial nephritis, acute renal failure, and nephrotic disorders. Non-limiting examples of prenatal diseases that can be inferred using the disclosed methods and systems include Down syndrome, aneuploidy, spina bifida, trisomy, Edwards syndrome, teratoma, sacrococcygeal teratoma (SCT), ventriculomegaly, renal agenesis, cystic fibrosis, and hydrops fetali. Non-limiting examples of metabolic diseases that can be inferred using the disclosed methods and systems include cystinosis, Fabry disease, Gaucher disease, Lesch-Nyhan syndrome, Niemann-Pick disease, phenylketonuria, Pompe disease, and Tay-Sachs disease.

[0326] Specific details of the particular examples may be combined in any suitable manner without departing from the spirit and scope of the disclosed examples of the inventive concept. However, other examples of the inventive concept may be directed to specific examples relating to each individual aspect or a particular combination of these individual aspects. All patents, patent applications, publications, and descriptions mentioned herein are incorporated herein by reference in their entirety for all purposes.

[0327] VI. Reagent Kit This disclosure provides a kit for identifying or monitoring cancer in a subject. The kit may include probes for identifying quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences at each of multiple cancer-related genomic loci in a cell-free biological sample of the subject. The quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences at each of multiple cancer-related genomic loci in the cell-free biological sample may indicate one or more types of cancer. The probes may selectively target sequences at multiple cancer-related genomic loci in the cell-free biological sample. The kit may include instructions for treating a cell-free biological sample with the probes to generate a dataset indicating quantitative measures (e.g., indicating presence, absence, or relative quantity) of sequences at each of multiple cancer-related genomic loci in the cell-free biological sample of the subject. In one embodiment, the kit includes a primer set, PCR reaction components, sequencing reagents, minimally destructive transformation reagents, and library preparation reagents.

[0328] The probes in the kit can selectively target sequences at multiple cancer-associated genomic loci in cell-free biological samples. The probes in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can have sequence complementarity with one or more nucleic acid sequences from multiple cancer-associated genomic loci or genomic regions. Multiple cancer-associated genomic loci or genomic regions can include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different cancer-associated genomic loci or genomic regions identified by targeted methylation sequencing. Multiple cancer-associated genomic loci or genomic regions may include 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 different cancer-associated genomic loci or genomic regions identified by targeted methylation sequencing.

[0329] The kit instructions may include instructions on using probes selectively targeting sequences at multiple cancer-associated genomic loci in a cell-free biological sample to determine the sample. These probes may be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity with one or more nucleic acid sequences (e.g., RNA or DNA) from multiple cancer-associated genomic loci. These nucleic acid molecules may be primers or enriched sequences. Instructions for determining the cell-free biological sample may include descriptions of performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample to generate a dataset indicating a quantitative measure (e.g., indicating presence, absence, or relative quantity) of the sequence at each of the multiple cancer-associated genomic loci in the cell-free biological sample. This quantitative measure (e.g., indicating presence, absence, or relative quantity) may indicate one or more cancers.

[0330] The kit's instructions may include instructions on measuring and interpreting assay readouts that can be quantified at one or more of multiple cancer-associated genomic loci to generate a dataset indicating a quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple cancer-associated genomic loci can generate a dataset indicating a quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple cancer-associated genomic loci in a cell-free biological sample. Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values ​​thereof.

[0331] Example Example 1 DNA extraction and library preparation using a transformation-tolerant sequencing adaptor / primer system targeting EM-seq.

[0332] A. DNA preparation Starting material: 2 mL plasma According to the manufacturer's protocol, DNA was purified using QIAamp circulating nucleic acid, extracted from 2 mL of plasma and eluted to a volume of 30 μL.

[0333] B. dsDNA denatures into ssDNA The extracted DNA sample (containing dsDNA) was denatured to generate ssDNA by incubating plates at 98°C for approximately 3 minutes on a thermal cycler. Immediately after removal from the thermal cycler, the denatured sample was placed on a frozen plate block on ice.

[0334] C. Transformation-tolerant adaptor connection In this embodiment, the functional sets of the transformation tolerance adaptor, PCR primers, and sequencing primers were tested as described in Example 2. Sequencing library yields were determined for libraries generated using the transformation tolerance adaptor or an adaptor containing 5mC.

[0335] Preparation of transformation-resistant single-stranded DNA library adaptors including two adaptor pairs (5' and 3', referred to as adaptor A and adaptor B, respectively). Figure 2The first adaptor pair consists of a top DNA oligonucleotide containing a 5' terminal amino group containing a 12-carbon spacer region. All cytosines in the top adaptor pair are 5 mC. The 3' sequence of each top adaptor pair contains one of several mCp-containing motifs for evaluating enzymatic oxidation and deamination performance. Each top A adaptor pair is annealed to a bottom adaptor pair containing unmethylated cytosine (adaptor A'). The bottom adaptor pair has a sequence that is completely complementary to the top adaptor pair, except for a random 7-nucleotide 5' overhang sequence. The bottom adaptor pair is modified with a 5' amino group containing a 6-carbon spacer region and a 3' amino group. The second adaptor pair consists of a top DNA oligonucleotide containing a 5' phosphate group, followed by one of several mCp-containing motifs, a sequence in which all cytosines are methylated, and a 3' dideoxycytosine. Each top B adaptor pair is annealed to a bottom adaptor pair containing unmethylated cytosine (adaptor B'). Apart from the random 7-nucleotide 3' overhang sequence, the bottom adaptor strand has a sequence that is completely complementary to the top adaptor strand. All lower B adaptor oligonucleotides are blocked by 5' terminal amino and 3' amino modifiers containing 12-carbon spacers. Alternatively, transformation-resistant adaptors can be generated by excluding all cytosine from the top adaptor sequence.

[0336] Approximately 20 μL of cfDNA input was dispensed into 96-well plates. The plates were sealed, vortexed, and centrifuged downwards. The samples were then placed in a thermal cycler at 98°C (105°C lid temperature) for 3 minutes. The samples were removed from the thermal cycler and immediately placed on an ice-cold plate stage for 5 minutes. Subsequently, 2 μL of each of the 5' and 3' ssDNA adapter pairs were added to the ligation mixture, and then added to each sample and mixed 10 times with a pipette. Next, 26 μL of ligation buffer was added to the samples and mixed 10 times with a pipette. The plates were sealed, centrifuged downwards, vortexed, and centrifuged downwards again. The plates were then incubated in a thermal cycler at 37°C (45°C lid temperature) for 1 hour.

[0337] After a 1-hour incubation, the plates were removed from the thermal cycler for bead-based purification. First, a diluted bead mixture was prepared: 75 μL of 10 mM Tris-HCl (pH 8.5) was added to 58.5 μL of pre-diluted buffer-displaced AMPure beads. Approximately 133.5 μL of this bead mixture was added to each sample. The sample-bead assemblies were then incubated on a workbench at room temperature for approximately 15 minutes. Next, the plates were placed on a Permagen bar magnet for 5 minutes or until the supernatant became visually clear. The supernatant was then removed without disturbing the magnetized bead clumps. While the plates were still on the magnetized rack, 200 μL of freshly prepared 80% ethanol (EtOH) was added to each well. After 30 seconds, the EtOH was removed. This washing process was then repeated (adding 200 μL of ethanol and removing it after 30 seconds). The plates were then removed from the magnet, sealed, and centrifuged using a pulsed centrifugation technique. Return the plate to the magnet and remove any remaining EtOH using a small pipette. Allow the magnetized beads to dry on the work surface at room temperature for 2 minutes, then remove them from the magnet. While demagnetized, add 16 μL of 10 mM Tris-HCl to the beads and mix thoroughly with a pipette. Incubate the beads at room temperature for 5 minutes, then return the plate to the magnet. When the supernatant becomes clear, transfer 15 μL of the supernatant to a new plate.

[0338] D. Second-chain synthesis In an alternative implementation, a second-strand synthesis (SSS) operation can be added after adaptor ligation and before the enzymatic oxidation reaction to convert the ssDNA library into a dsDNA library. After ssDNA library preparation, bead purification is performed. Then, KAPA PCR mixture, Tris-HCl, and 2 μL of 50 μM primers (complementary to the 3' adaptor) are added to the library. The sample is vortexed and then subjected to a single extension reaction on a thermal cycler. After this extension, the sample is washed with purification beads and eluted in elution buffer.

[0339] E. Unpurified second-chain synthesis No purification was performed after ssDNA library preparation. Immediately following adaptor ligation, the following were added to each sample: 1 μL of 50 μM primer complementary to the 3' adaptor, 0.5 μL of dNTPs (from the Fast Start kit), 0.2 μL of LFast Taq polymerase, and water to increase volume. The samples were vortexed and then extended by centrifugation on a thermal cycler. After extension, the samples were washed with purification beads and eluted in elution buffer.

[0340] F. Unpurified second-strand synthesis using alternative linkers Single-strand ligation was performed as described above; however, the bottom adaptor B sequence was truncated, resulting in incomplete complementarity with the top adaptor B sequence. No purification was performed after ssDNA library preparation. Immediately following adaptor ligation, the following were added to each sample: 1 μL of 50 μM primer complementary to the 3' adaptor, 0.5 μL of dNTPs (from the Fast Start kit), 0.2 μL of LBst polymerase, and water to increase volume. The samples were vortexed and centrifuged, then subjected to one PCR cycle on a thermal cycler. After extension, the samples were washed with purification beads and eluted in elution buffer.

[0341] Subsequently, ssDNA libraries with transformation-tolerant adaptors (whether single-stranded or double-stranded (with or without second-strand synthesis)) were used as starting materials for the subsequent methylation transformation and sequencing reactions described herein in Example 2.

[0342] Example 2 Preparation of targeted EM-seq libraries.

[0343] A. Oxidation of 5-methylcytosine and 5-hydroxymethylcytosine Prepare the TET2 reaction buffer according to the manufacturer's instructions. Then, add the TET2 reaction buffer to a tube of TET2 reaction buffer supplement and mix thoroughly. On ice, add the TET2 reaction buffer, oxidation supplement, oxidation enhancer, and TET2 enzyme directly to the ssDNA library (“sample”) prepared according to Example 1. Then, thoroughly mix each sample mixture by vortexing. After a brief centrifugation of the mixture, add the iron solution to the mixture. Then, thoroughly mix the mixture by vortexing or by pipetting up and down, and briefly centrifuge. Then, incubate the mixture in a thermal cycler at 37°C for 1 hour. Then, transfer the mixture to ice, treat with 1 μl of the stop reagent, and the mixture should appear yellow. Then, thoroughly mix the mixture by vortexing or by pipetting up and down at least 10 times, and briefly centrifuge. Finally, incubate the mixture in a thermal cycler at 37°C for 30 minutes, then at 4°C.

[0344] B. Purify the DNA transformed by TET2 NEBNext Vortex the sample purification beads and add them to each sample, then mix thoroughly by up-and-down pipetting. Incubate the sample on the workbench at room temperature for at least 5 minutes. Then, place the tube against a suitable magnetic rack to separate the beads from the supernatant. After 5 minutes (or when the solution is clear), carefully remove the supernatant without disturbing the DNA-targeting beads and discard it. While on the magnetic rack, add freshly prepared 80% ethanol to each tube. Incubate the sample at room temperature for 30 seconds, then carefully remove and discard the supernatant. Repeat the wash once, for a total of two washes. After the second wash, remove all visible liquid using a p10 pipette tip. Then allow the beads to air dry for 2 minutes while the tubes are left uncapped and placed on the magnetic rack. Then, remove the tubes from the magnetic rack. Elute the DNA from the beads with elution buffer. Add the elution buffer to each tube and mix thoroughly by up-and-down pipetting 10 times. Then, incubate the sample at room temperature for at least 1 minute. If necessary, rapidly centrifuge the sample to collect liquid from the side of the tube before placing it back on the magnetic rack. Then place the tube back on the magnetic rack. After 3 minutes (or at any time when the solution is clear), transfer the eluted DNA from the supernatant to a new PCR tube.

[0345] C. DNA denaturation DNA samples (containing dsDNA) were denatured by incubating plates at 85°C for 10 minutes on a thermal cycler. Immediately after removal from the thermal cycler, the DNA samples were placed on a plate freezing stage on ice.

[0346] D. Deamination of cytosine Add APOBEC reaction buffer, bovine serum albumin (BSA), and APOBEC to the denatured DNA. The mixture is then thoroughly mixed by vortexing or by pipetting up and down at least 10 times, followed by a brief centrifugation. The mixture is then incubated according to the following protocol: at 4°C for 10 minutes, then increasing the temperature by 1°C every 2 minutes and 15 seconds for 46 cycles until the sample reaches 50°C. The mixture is then held at 50°C for 10 minutes in a thermal cycler, followed by maintenance at 4°C.

[0347] E. Thermosensitive proteinase K treatment To halt APOBEC activity, heat-sensitive proteinase K (TLPK) was added to the sample after deamination and incubated at 37°C for 30 minutes, followed by inactivation at 65°C for 10 minutes. After the first TLPK treatment, the sample was amplified by PCR using index primers. Amplification could be run any number of times. After PCR, a second TLPK treatment was performed. Following the second TLPK treatment, bead purification was performed.

[0348] F. Purification of deamination-treated DNA NEBNext Vortex the sample purification beads and add them to each sample, then mix thoroughly by pipetting up and down at least 10 times. During the final mix, carefully remove all liquid from the pipette tip. Then, incubate the sample on the workbench at room temperature for at least 5 minutes. After 5 minutes (or when the solution is clear), carefully remove and discard the supernatant. While on the magnetic rack, add freshly prepared 80% ethanol to the tube. Then, incubate the sample at room temperature for 30 seconds, then carefully remove and discard the supernatant. Repeat the wash once, for a total of two washes. Next, allow the beads to air dry for 90 seconds while the tube is left uncapped and placed on the magnetic rack. Then, elute the DNA target from the beads with elution buffer. Add the elution buffer to each tube and mix thoroughly by pipetting up and down 10 times. Incubate the sample at room temperature for at least 1 minute. If necessary, rapidly centrifuge the sample before returning the tube to the magnetic rack to collect liquid from the side of the tube. Then, return the tube to the magnetic rack. After 3 minutes (or at any time when the solution becomes clear), transfer the eluted DNA target from the supernatant to a new PCR tube.

[0349] Example 3 Target capture and multiplex amplification.

[0350] After quantification, the samples were pooled and concentrated. The enzymatically transformed library was then target-enriched to specifically enrich pre-identified DNA fragments containing target CpG sites using 5' biotinylated capture probes. These probes could be methylated, unmethylated, or a combination thereof. Hybrid selection was performed using the TWIST Rapid Hybridization Target Capture Kit. Hybridization could occur at 58°C, 60°C, or other temperatures. Bead washing was performed at a heated temperature (e.g., hybridization temperature + 3°C). After hybridization, the captured DNA fragments were amplified by PCR. The target capture library was sequenced on an Illumina Novaseq sequencer using 2 x 150 cycle runs.

[0351] Example 4 Targeted methylation classification.

[0352] The raw data files are used for alignment and methylation interpretation to allow for targeted methylation analysis of pre-identified regions of the genome. Whole-genome amplification of the enzymatically transformed DNA is then performed.

[0353] FASTQ files were mapped to a reference genome, and methylation scores were calculated for disease classification. Characterized data, including sets of CpG sites associated with health, disease, disease state, and treatment responsiveness, were processed using a machine learning model to identify a classifier that stratifies individuals in the population based on high-methylation model scores and / or low-methylation model scores.

[0354] Example 5 Comparison of ssDNA and dsDNA library preparation in methylation studies A. Advanced adenoma (AA) and colorectal cancer (CRC) Blood samples were obtained from 11 healthy human subjects, 23 human subjects with advanced adenomas (AA), and 30 human subjects with colorectal cancer (CRC). DNA was extracted from the blood samples and libraries were prepared according to the following: (i) the ssDNA library preparation method described in Example 1 and (ii) other dsDNA library preparation methods to generate ssDNA and dsDNA libraries, respectively (E7120 NEBNext). Enzyme-catalyzed methyl-seq kit E7120).

[0355] Both the ssDNA library (generated using the methods disclosed herein) and the dsDNA library underwent targeted methylation transformation sequencing as described in Examples 2-3. The raw data files were used for alignment and methylation interpretation to allow targeted methylation analysis of pre-identified regions of the genome against both high and low methylation model scores.

[0356] B. Result Figure 5 A comparison is provided of hypermethylation model scores for cfDNA extracted from healthy / cancer-negative (NEG), advanced adenoma (AA), and colorectal cancer (CRC) samples, detected using the ssDNA library preparation method described in Example 1 and another dsDNA library preparation method. When processed using both ssDNA library preparation and another dsDNA library preparation method, the hypermethylation rate of cfDNA derived from selected genomic regions distinguished cfDNA derived from AA and CRC patients from cancer-negative cfDNA. This experiment demonstrates that the ssDNA library preparation method described herein can produce methylation analysis (e.g., hypermethylation scores) equivalent to other dsDNA library preparation methods. Therefore, for example, when performing DNA hypermethylation analysis, this workflow is equivalent to an end-repair-dependent method in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0357] Figure 7A comparison is provided of the hypomethylation scores of cfDNA extracted from healthy / cancer-negative (NEG), advanced adenoma (AA), and colorectal cancer (CRC) samples, detected using the ssDNA library preparation method described in Example 1 and another dsDNA library preparation method. When the ssDNA library preparation method described herein is used instead of the dsDNA library preparation method, the hypomethylation rate of cfDNA derived from selected genomic regions distinguishes AA and CRC patient-derived cfDNA from cancer-negative cfDNA. Therefore, this example demonstrates that the ssDNA library preparation method described herein optimizes recovery, methylation analysis, quantification, and target capture. Thus, in some embodiments, such as when using DNA hypomethylation analysis, this workflow can outperform end-repair-dependent methods in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0358] C. Liver cancer, lung cancer, and pancreatic cancer Blood samples were obtained from 11 healthy human subjects, 10 human subjects with liver cancer, 27 human subjects with lung cancer, and 21 human subjects with pancreatic cancer. DNA was extracted from the blood samples and libraries were prepared according to the following: (i) the ssDNA library preparation method described in Example 1 and (ii) other dsDNA library preparation methods (E7120 NEBNext). Enzyme-catalyzed methyl-seq kit E7120).

[0359] Both the ssDNA library (generated using the methods disclosed herein) and the dsDNA library underwent targeted methylation transformation sequencing as described in Examples 2-3. The raw data files were used for alignment and methylation interpretation to allow targeted methylation analysis of pre-identified regions of the genome against both high and low methylation model scores.

[0360] D. Result Figure 6A comparison is provided of hypermethylation model scores for cfDNA extracted from healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) samples, detected using the ssDNA library preparation method described herein and other dsDNA library preparation methods. When processed using the ssDNA library preparation method described herein and other dsDNA library preparation methods, the hypermethylation rate of cfDNA derived from selected genomic regions distinguishes cfDNA derived from hepatocellular carcinoma, lung cancer, and pancreatic cancer patients from cancer-negative cfDNA. This example demonstrates that the ssDNA library preparation method described herein produces methylation analysis (e.g., hypermethylation scores) equivalent to other dsDNA library preparation methods. Therefore, for example, when using DNA hypermethylation analysis, the ssDNA library preparation workflow described herein is equivalent to an end-repair-dependent method in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0361] Figure 8 This paper provides a comparison of hypomethylation scores of cfDNA extracted from samples of healthy / cancer-negative (NEG), hepatocellular carcinoma (liver), lung cancer (lung), and pancreatic cancer (pancreas) patients, detected using the ssDNA library preparation method and the dsDNA library preparation method described herein. Figure 8 As shown, when using ssDNA library preparation instead of dsDNA library preparation, the hypomethylation rate of cfDNA derived from selected genomic regions distinguishes cfDNA derived from liver, lung, and pancreatic cancer patients from cancer-negative cfDNA. This example provides additional evidence that the ssDNA library preparation method described herein optimizes recovery, methylation analysis, quantification, and target capture across multiple cancer types, thereby further clarifying that this workflow can outperform end-repair-dependent methods, for example, when using DNA hypomethylation analysis, in distinguishing cfDNA samples derived from cancer patients from those derived from healthy patients.

[0362] While preferred embodiments of the inventive concept have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The inventive concept is not intended to be limited to the specific examples provided herein.

[0363] While the inventive concept has been described with reference to the foregoing specification, the description and illustration of embodiments herein are not intended to be construed in a limiting sense. Many variations, modifications, and alternatives will now occur to those skilled in the art without departing from the inventive concept. Furthermore, it should be understood that all aspects of the inventive concept are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables.

[0364] It should be understood that various alternatives to the embodiments of the inventive concept described herein may be employed in the practice of the inventive concept. Therefore, the inventive concept is intended to also cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the inventive concept and thereby cover the methods and structures within the scope of these claims and their equivalents.

Claims

1. A method for preparing a sequencing library for methylation sequencing of one or more nucleic acid molecules of a biological sample or its derivatives, comprising: (a) Obtaining a nucleic acid composition, wherein the nucleic acid composition comprises a plurality of single-stranded nucleic acid molecules obtained or derived from the biological sample; (b) Linking a nucleic acid adaptor to one of the plurality of single-stranded nucleic acid molecules to generate an adaptor-linked nucleic acid molecule, wherein the nucleic acid adaptor comprises a nucleic acid resistant to base conversion by a methylation enrichment method; and (c) subjecting the adaptor-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed adaptor-linked nucleic acid molecule.

2. The method of claim 1, wherein the ligation in (b) further comprises treatment with a deoxyribonucleic acid (DNA) ligase.

3. The method of claim 1, wherein the connection in (b) further comprises treatment with a polynucleotide kinase.

4. The method of claim 1, wherein the nucleic acid adaptor comprises a double-stranded oligonucleotide comprising an adaptor sequence and a protruding end sequence.

5. The method of claim 4, wherein the protruding end sequence is a 3' protruding end sequence.

6. The method of claim 4, wherein the overhang sequence comprises a random sequence of oligonucleotides.

7. The method of claim 1, wherein the nucleic acid adaptor comprises one or more methylated cytosine bases.

8. The method according to claim 1, wherein the nucleic acid adaptor does not contain methylated or unmethylated cytosine bases.

9. The method of claim 1, wherein the nucleic acid adaptor comprises a unique molecular identifier.

10. The method of claim 9, wherein the unique molecular identifier is configured to measure the enrichment efficiency of the methylation conversion method.

11. The method of claim 1, further comprising treating the biological sample or a derivative thereof prior to (a) to generate the plurality of single-stranded nucleic acid molecules.

12. The method of claim 11, wherein the treatment comprises denaturing double-stranded nucleic acid molecules in the biological sample or a derivative thereof.

13. The method of claim 12, wherein the denaturation further comprises applying heat to the double-stranded nucleic acid molecule.

14. The method of claim 12, wherein the denaturation further comprises applying heat to the double-stranded nucleic acid molecule and then rapidly cooling the denatured single-stranded nucleic acid molecule.

15. The method of claim 1, further comprising treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex.

16. The method of claim 1, wherein the method does not include treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex.

17. The method according to claim 15 or 16, wherein the binding agent is a single-stranded nucleic acid binding protein (SSB).

18. The method of claim 1, further comprising amplifying the transformed adaptor-linked nucleic acid molecule.

19. The method of claim 18, wherein the amplification comprises polymerase chain reaction (PCR).

20. The method of claim 1, further comprising contacting the transformed adaptor-linked nucleic acid molecule or a derivative thereof with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence at least partially complementary to the CpG or CH locus of the reference group.

21. The method of claim 20, wherein the nucleic acid probe comprises an unmethylated nucleic acid probe.

22. The method of claim 20, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at the CpG locus in the CpG or CH locus from the reference group.

23. The method of claim 20, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at the CpG locus in the CpG or CH locus from the reference group.

24. The method according to any one of claims 20-23, further comprising determining the nucleic acid sequence of the enriched nucleic acid molecule or a derivative thereof.

25. The method according to any one of claims 20-24, further comprising sequencing the enriched nucleic acid molecules or derivatives thereof to generate sequencing data.

26. The method of claim 25, further comprising analyzing the sequencing data to generate a methylation profile of the nucleic acid molecule of the biological sample or a derivative thereof.

27. The method of claim 26, wherein the analysis further comprises comparing the sequencing data with a reference sequence.

28. The method of claim 1, further comprising subjecting the nucleic acid molecule linked by the adaptor to an extension reaction after (b) to generate a partially double-stranded nucleic acid molecule or a fully double-stranded nucleic acid molecule.

29. The method of claim 28, wherein the extension reaction is carried out in the presence of polymerase, a variety of deoxynucleoside triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

30. The method according to claim 1, wherein the nucleic acid molecule is deoxyribonucleic acid (DNA).

31. The method of claim 30, wherein the DNA is cell-free DNA.

32. The method according to claim 1, wherein the biological sample is a cell-free biological sample.

33. The method of claim 32, wherein the cell-free biological sample is a plasma sample.

34. The method of claim 1, wherein the methylation enrichment method comprises treatment with one or more enzymes.

35. The method of claim 1, wherein the methylation enrichment method comprises treatment with a deca-echotranslocation (TET) enzyme.

36. The method of claim 1, wherein the methylation enrichment method does not include treatment with bisulfite.

37. A method for preparing a sequencing library for methylation sequencing of nucleic acid molecules from biological samples or their derivatives, comprising: (a) Obtaining a nucleic acid composition, wherein the nucleic acid composition comprises a plurality of single-stranded nucleic acid molecules obtained or derived from the biological sample; (b) subjecting the single-stranded nucleic acid molecules among the plurality of single-stranded nucleic acid molecules to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating transformed single-stranded nucleic acid molecules; and (c) Linking a nucleic acid adaptor to the transformed single-stranded nucleic acid molecule to generate an adaptor-linked transformed nucleic acid molecule, wherein the nucleic acid adaptor contains a nucleic acid resistant to base transformation by the methylation enrichment method.

38. The method of claim 37, wherein the ligation in (c) further comprises treatment with a deoxyribonucleic acid (DNA) ligase.

39. The method of claim 37, wherein the connection in (c) further comprises treatment with a polynucleotide kinase.

40. The method of claim 37, wherein the nucleic acid adaptor comprises a double-stranded oligonucleotide comprising an adaptor sequence and a protruding end sequence.

41. The method of claim 40, wherein the protruding end sequence is a 3' protruding end sequence.

42. The method of claim 40, wherein the overhang sequence comprises a random sequence of oligonucleotides.

43. The method of claim 37, wherein the nucleic acid adaptor comprises one or more methylated cytosine bases.

44. The method of claim 37, wherein the nucleic acid adaptor does not contain methylated or unmethylated cytosine bases.

45. The method of claim 37, wherein the nucleic acid adaptor comprises a unique molecular identifier.

46. ​​The method of claim 45, wherein the unique molecular identifier is configured to enable the measurement of the enrichment efficiency of the methylation enrichment method.

47. The method of claim 37, further comprising treating the biological sample or a derivative thereof prior to (a) to generate the plurality of single-stranded nucleic acid molecules.

48. The method of claim 47, wherein the treatment comprises denaturing double-stranded nucleic acid molecules in the biological sample or a derivative thereof.

49. The method of claim 47, wherein the denaturation comprises applying heat to the double-stranded nucleic acid molecule.

50. The method of claim 47, wherein the denaturation comprises applying heat to the double-stranded nucleic acid molecules to produce the plurality of single-stranded nucleic acid molecules, followed by rapid cooling of the plurality of single-stranded nucleic acid molecules.

51. The method of claim 37, further comprising treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex.

52. The method of claim 37, wherein the method does not include treating the single-stranded nucleic acid molecule or a derivative thereof with a binding agent configured to reduce the likelihood of forming a nucleic acid duplex.

53. The method according to claim 51 or 52, wherein the binding agent is a single-stranded nucleic acid binding protein (SSB).

54. The method of claim 37, further comprising amplifying the transformed nucleic acid molecule linked by the adaptor.

55. The method of claim 54, wherein the amplification comprises polymerase chain reaction (PCR).

56. The method of claim 37, further comprising contacting the adapted nucleic acid molecule or a derivative thereof linked by the adaptor with a nucleic acid probe to generate an enriched nucleic acid molecule, wherein the nucleic acid probe comprises a nucleic acid sequence at least partially complementary to the CpG or CH locus of the reference group.

57. The method of claim 56, wherein the nucleic acid probe comprises an unmethylated nucleic acid probe.

58. The method of claim 56, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at the CpG locus in the CpG or CH locus from the reference group.

59. The method of claim 56, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at the CpG or CH locus from the reference group.

60. The method according to any one of claims 56-59, further comprising determining the nucleic acid sequence of the enriched nucleic acid molecule or a derivative thereof.

61. The method according to any one of claims 56-59, further comprising sequencing the enriched nucleic acid molecules or derivatives thereof to generate sequencing data.

62. The method of claim 61, further comprising analyzing the sequencing data to generate a methylation profile of the nucleic acid molecule of the biological sample or a derivative thereof.

63. The method of claim 62, wherein the analysis further comprises comparing the sequencing data with a reference sequence.

64. The method of claim 37, further comprising subjecting the transformed nucleic acid molecule linked by the adaptor to an extension reaction after (b) to generate a partially double-stranded nucleic acid molecule or a fully double-stranded nucleic acid molecule.

65. The method of claim 64, wherein the extension reaction is carried out in the presence of polymerase, a variety of deoxynucleoside triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

66. The method of claim 37, wherein the nucleic acid molecule is deoxyribonucleic acid (DNA).

67. The method of claim 66, wherein the DNA is cell-free DNA.

68. The method of claim 37, wherein the biological sample is a cell-free biological sample.

69. The method of claim 68, wherein the cell-free biological sample is a plasma sample.

70. The method of claim 37, wherein the methylation enrichment method comprises treatment with one or more enzymes.

71. The method of claim 37, wherein the methylation enrichment method comprises treatment with a deca-echotranslocation (TET) enzyme.

72. The method of claim 37, wherein the methylation enrichment method does not include treatment with bisulfite.

73. A method comprising: (a) Connecting a nucleic acid intransitive linker to a single-stranded nucleic acid molecule obtained or derived from a biological sample of the object to generate an intransitive linker-connected nucleic acid molecule, wherein the nucleic acid intransitive linker comprises a nucleic acid resistant to base conversion by a methylation enrichment method; (b) subjecting the adaptor-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed adaptor-linked nucleic acid molecule; (c) Amplify the nucleic acid molecule linked by the transformed adaptor to generate the amplified nucleic acid molecule; (d) Contact the amplified nucleic acid molecule or its derivative with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence that is at least partially complementary to the CpG or CH locus of the reference group. (e) Determine the nucleic acid sequence of the enriched nucleic acid molecule or its derivative; (f) Compare the nucleic acid sequence of the enriched nucleic acid molecule or its derivative with a reference nucleic acid sequence; as well as (g) Training a machine learning model to generate a classifier that distinguishes between objects with cancer and objects without said cancer, wherein said machine learning model is trained with methylation profiles generated from: (i) a first set of nucleic acid samples from objects with said cancer; and (ii) a second set of nucleic acid samples from objects without said cancer.

74. The method of claim 73, wherein the amplification comprises polymerase chain reaction (PCR).

75. The method of claim 73, further comprising determining the nucleic acid sequence of the enriched nucleic acid molecule or a derivative thereof at a depth greater than 100x.

76. The method of claim 73, wherein the nucleic acid probe comprises unmethylated nucleic acid.

77. The method of claim 73, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to unmethylated cytosine bases at the CpG locus in the CpG or CH locus from the reference group.

78. The method of claim 73, wherein the nucleic acid probe is configured to selectively hybridize with one or more target regions of interest, the one or more target regions of interest corresponding to methylated cytosine bases at the CpG locus in the CpG or CH locus from the reference group.

79. The method according to any one of claims 73-78, further comprising sequencing the enriched nucleic acid molecules or derivatives thereof to generate sequencing data.

80. The method of claim 79, further comprising analyzing the sequencing data to generate a methylation profile of the nucleic acid molecule of the biological sample or a derivative thereof.

81. The method of claim 73, further comprising subjecting the transformed nucleic acid molecule linked by the adaptor after (a) to an extension reaction to generate a partially double-stranded nucleic acid molecule or a fully double-stranded nucleic acid molecule.

82. The method of claim 81, wherein the extension reaction is carried out in the presence of polymerase, a variety of deoxyribonucleoside triphosphates (dNTPs), and a primer complementary to the 3' end of the nucleic acid adaptor.

83. The method of claim 73, wherein the single-stranded nucleic acid molecule is deoxyribonucleic acid (DNA).

84. The method of claim 83, wherein the DNA is cell-free DNA.

85. The method of claim 73, wherein the biological sample is a cell-free biological sample.

86. The method of claim 85, wherein the cell-free biological sample is a plasma sample.

87. The method of claim 73, wherein the methylation enrichment method comprises treatment with one or more enzymes.

88. The method of claim 73, wherein the methylation enrichment method comprises treatment with a deca-eicosyltransferase (TET).

89. The method of claim 73, wherein the methylation enrichment method does not include treatment with bisulfite.

90. The method of claim 73, wherein the reference set comprises a CpG or CH locus associated with a transcription start site.

91. The method of claim 73, further comprising identifying the tissue of origin of the nucleic acid molecule.

92. The method of claim 73, further comprising identifying the genomic location and fragment length of the nucleic acid molecule.

93. The method of claim 73, wherein the machine learning model is trained using feature inputs selected from the following: base-wise methylation percentage of CpG, base-wise methylation percentage of CHG, base-wise methylation percentage of CHH, count or ratio of methylated CpG fragments with different counts or ratios in the observed region, conversion efficiency, low-methylated blocks, methylation level of CpG, methylation level of CHH, methylation level of CHG, fragment length, fragment midpoint, methylation level of chrM, methylation level of LINE1, methylation level of ALU, dinucleotide coverage, uniformity of coverage, global average CpG coverage, average coverage at CpG islands, CGI racks, and CGI shores.

94. The method of claim 73, wherein the classifier distinguishing between subjects suffering from said cancer and subjects not suffering from said cancer comprises: A set of measurements representing methylation profiles from methylation sequencing data derived from subjects with and without the cancer. The measured values ​​are used to generate a feature set corresponding to the characteristics of the methylation spectrum, wherein the feature set is processed by a machine learning or statistical model, wherein the machine learning or statistical model provides a feature vector that can be used as a classifier to distinguish between a group of objects with the cancer and a group of objects without the cancer.

95. The method of claim 73, further comprising: (a) The methylation profile of the biological sample was obtained by determination using the methylation enrichment method; (b) Classifying the methylation spectrum of the biological sample as an indication of the presence of the cancer in the object using a trained machine learning algorithm; as well as (c) If the trained machine learning algorithm classifies the biological sample as negative for the cancer at a specified confidence level, then outputs a report identifying the biological sample as negative for the cancer.

96. The method of claim 73, further comprising: (a) Determine the baseline methylation profile of the biological sample of the object in its baseline methylation state; (b) Determine the test methylation profiles of biological samples of the object at one or more time points after the baseline methylation state; as well as (c) Determine the change in the test methylation spectrum compared to the baseline methylation spectrum, wherein the change indicates a change in the minor residual lesion state of the object.

97. The method of claim 96, wherein the minimal residual disease state is selected from: treatment response, tumor burden, postoperative tumor residue, recurrence, secondary screening, primary screening, and cancer progression.

98. The method according to claim 95 or 96, wherein the methylation profile includes hypermethylation analysis and / or hypomethylation analysis.

99. The method of claim 73, wherein the cancer comprises two or more of colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer.

100. The method of claim 73, wherein the cancer is colorectal cancer, breast cancer, pancreatic cancer, liver cancer, or lung cancer.

101. The method of claim 100, wherein the cancer is the colorectal cancer.

102. The method of claim 100, wherein the cancer is the lung cancer.

103. The method of claim 100, wherein the cancer is the liver cancer.

104. The method of claim 100, wherein the cancer is the pancreatic cancer.

105. A method comprising: (a) Connecting a nucleic acid intransitive linker to a single-stranded nucleic acid molecule obtained or derived from a biological sample of the object to generate an intransitive linker-connected nucleic acid molecule, wherein the nucleic acid intransitive linker comprises a nucleic acid resistant to base conversion by a methylation enrichment method; (b) subjecting the adaptor-linked nucleic acid molecule to conditions sufficient to convert unmethylated cytosine to uracil using a methylation enrichment method, thereby generating a transformed adaptor-linked nucleic acid molecule; (c) Amplify the nucleic acid molecule linked by the transformed adaptor to generate the amplified nucleic acid molecule; (d) Contact the amplified nucleic acid molecule or its derivative with a nucleic acid probe to generate enriched nucleic acid molecules, wherein the nucleic acid probe comprises a nucleic acid sequence that is at least partially complementary to the CpG or CH locus of the reference group. (e) Determine the nucleic acid sequence of the enriched nucleic acid molecule or its derivative; (f) Compare the nucleic acid sequence of the enriched nucleic acid molecule or its derivative with a reference nucleic acid sequence; as well as (g) Training a machine learning model to generate a classifier that distinguishes between objects with the indication and objects without the indication, wherein the machine learning model is trained with methylation profiles generated from: (i) a first set of nucleic acid samples from objects with the indication; and (ii) a second set of nucleic acid samples from objects without the indication.

106. The method of claim 105, further comprising: (h) The methylation profile of the biological sample is obtained by determination using the methylation enrichment method; (i) Classify the methylation spectrum of the biological sample as an indication of the presence of the indication in the object using a trained machine learning algorithm; as well as (j) If the trained machine learning algorithm classifies the biological sample as negative for the indication at a specified confidence level, then outputs a report identifying the biological sample as negative for the indication.

107. The method of claim 105, further comprising: (h) Determine the baseline methylation profile of the biological sample of the object in the baseline methylation state; (i) Determine the test methylation profiles of biological samples of the object at one or more time points after the baseline methylation state; as well as (j) Determine the change in the tested methylation spectrum compared to the baseline methylation spectrum, wherein the change indicates a change in the minimal residual disease status of the indication in the subject.

108. The method of claim 107, wherein the minimal residual disease state is selected from: treatment response, recurrence, secondary screening, primary screening, and indication progression.

109. The method according to claim 106 or 107, wherein the methylation profile includes hypermethylation analysis and / or hypomethylation analysis.

110. The method of claim 105, wherein the indications include bowel-related diseases, immune-mediated inflammatory diseases, nervous system diseases, kidney diseases, prenatal diseases, or metabolic diseases.

111. A system comprising: (a) A computer-readable medium product comprising the classifier according to claim 73. The classifier includes: a set of measurements representing methylation spectra from methylation sequencing data, said methylation sequencing data being from subjects with and without said cancer; wherein said measurement set is used to generate a feature set corresponding to characteristics of the methylation spectra from said subjects with and without said cancer; wherein said feature set is processed by a machine learning or statistical model, said machine learning or statistical model providing feature vectors that can be used to distinguish between subjects with and without said cancer; and (b) One or more processors for executing instructions stored on the computer-readable medium product.

112. The system according to claim 111, wherein the classifier is selected from linear discriminant analysis (LDA) classifier, quadratic discriminant analysis (QDA) classifier, support vector machine (SVM) classifier, random forest (RF) classifier, linear kernel support vector machine classifier, first-order polynomial kernel support vector machine classifier, second-order polynomial kernel support vector machine classifier, ridge regression classifier, elastic network algorithm classifier, sequence minimum optimization algorithm classifier, naive Bayes algorithm classifier, and nonnegative matrix factorization (NMF) predictor algorithm classifier.

113. The system of claim 111 is further configured to perform any one or more of the methods described.

114. The system of claim 111, wherein the system includes one or more processors configured to perform any of the more than one of the methods.

115. The system of claim 111, wherein the system comprises modules that respectively perform the operations of any one or more of the methods.

116. A kit for detecting cancer, comprising reagents for performing any of the methods described above, and instructions for detecting cancer signals.

117. The kit according to claim 116, wherein the reagent is selected from primer sets, PCR reaction components, sequencing reagents, methylation enrichment reagents, and library preparation reagents.