Machine learning techniques for determining base methylation

Through machine learning models combined with the kinetic signals of single-molecule sequencing, DNA methylation is directly detected, which solves the degradation and CG bias of bisulfite sequencing methods, and realizes accurate measurement and efficient detection of DNA methylation.

CN120283284APending Publication Date: 2025-07-08CENT FOR NOVOSTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380080946.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-16
Filing Date
2023-12-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing bisulfite sequencing methods have problems in degrading DNA, CG bias and reduced signal-to-noise ratio when detecting DNA methylation, and are not suitable for sequencing of long DNA molecules, so they cannot accurately measure DNA methylation.

Method used

Using machine learning models, combined with convolutional neural networks and transformer layers, the kinetic signals of single-molecule sequencing are used to directly detect DNA methylation by capturing local and global signal patterns, avoiding enzymatic or chemical conversion steps, and improving the accuracy and efficiency of detection.

Benefits of technology

Accurate measurement of DNA methylation is achieved, the accuracy and efficiency of detection is improved, and the methylation pattern near the end of the DNA fragment is able to be identified without enzymatic or chemical transformation, which is suitable for a variety of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283284A_ABST
    Figure CN120283284A_ABST
Patent Text Reader

Abstract

Systems and methods for determining base methylation in analyzing nucleic acid molecules. Embodiments may use kinetic signals generated by DNA polymerases during single molecule sequencing. The use of the kinetic signal of the aptamer sequence may allow for the determination of the methylation pattern near the end of the DNA fragment. A deep learning model capable of capturing local and global signal patterns may be trained to detect base methylation using these features with improved performance. The deep learning model may include a convolutional neural network that preferentially captures features of a local signal pattern, which integrates a converter model that preferentially captures features of a global signal pattern. Improved performance in the assay of base methylation may result in a more accurate diagnosis of a subject. Accurate measurement of DNA methylation may have several other clinical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 433,253, filed on December 16, 2022, the entire content of which is incorporated herein by reference for all purposes. Background of the Invention

[0002] DNA methylation is an epigenetic mechanism by which a methyl group is added to a DNA base. For example, a methyl group can be covalently added to the 5th position of cytosine to form 5 - methylcytosine. Methylation has been found on cytosine, adenine, thymine, and guanine, such as 5mC (5 - methylcytosine), 6mA (N6 - methyladenine), 4mC (N4 - methylcytosine), 5hmC (5 - hydroxymethylcytosine), 5fC (5 - formylcytosine), 5caC (5 - carboxylcytosine), 1mA (N1 - methyladenine), 3mA (N3 - methyladenine), 7mA (N7 - methyladenine), 3mC (N3 - methylcytosine), 2mG (N2 - methylguanine), 6mg (O6 - methylguanine), 7mG (N7 - methylguanine), 3mT (N3 - methylthymine), and 4mT (O4 - methylthymine).

[0003] DNA methylation plays important biological roles, such as silencing retroviral elements, regulating tissue - specific gene expression, genomic imprinting, X - chromosome inactivation, tumorigenesis, and regulating many other diseases (Moore et al., Neuropsychopharmacology. 2013; 38:23 - 38). DNA methylation occurring in different genomic regions can have different effects on gene activity based on the underlying gene sequence. Therefore, precise measurement of DNA methylation will have many clinical applications.

[0004] Precise measurement of methylomic modifications on DNA molecules will have many clinical applications. A widely used method for measuring DNA methylation is through the use of bisulfite sequencing (BS - seq) (Lister et al., 2009; Frommer et al., 1992). In this method, a DNA sample is first treated with bisulfite, which converts unmethylated cytosine (i.e., C) to uracil. In contrast, methylated cytosine remains unchanged. Then the bisulfite - modified DNA is analyzed by DNA sequencing. In another method, after bisulfite conversion, polymerase chain reaction (PCR) amplification of the modified DNA is performed using primers capable of distinguishing bisulfite - converted DNA with different methylation profiles (Herman et al., 1996). The latter method is called methylation - specific PCR.

[0005] One drawback of this bisulfite-based method is that the bisulfite conversion step reportedly significantly degrades most of the DNA being processed (Grunau, 2001). Another drawback is that the bisulfite conversion step will produce a strong CG bias (Olova et al., 2018), often resulting in a reduced signal-to-noise ratio for DNA mixtures with heterogeneous methylation states. Additionally, due to the degradation of DNA during bisulfite treatment, bisulfite sequencing is not an ideal method for sequencing long DNA molecules.

[0006] The systems and methods described herein improve the accuracy and efficiency of detecting DNA methylation. The systems and methods can avoid the chemical conversion step for detecting methylation. The systems and methods can also be cheaper than other systems and methods. SUMMARY OF THE INVENTION

[0007] In the present disclosure, enhanced systems and methods determine base methylation in analyzing nucleic acid molecules. Embodiments can use kinetic signals generated by DNA polymerase during single molecule sequencing. The methods described herein can include using features derived from the kinetic signals of sequencing. These features can include the pulse width of the optical or electrical signal from the sequenced base, the inter-pulse duration of the base, and the identity of the base, which can be generated by a target DNA molecule extracted from an organism (e.g., a human) and an artificial sequence (aptamer) linked to the target DNA molecule. The use of the kinetic signals of the aptamer sequence can allow determination of the methylation pattern near the end of a DNA fragment, which may initially not be analyzable due to insufficient kinetic signals flanking the locus near the end of the fragment.

[0008] A machine learning model capable of capturing local and global signal patterns can be trained to use these features to detect base methylation with improved performance. The machine learning model can include a convolutional neural network that preferentially captures features of local signal patterns, which integrates a transformer model that preferentially captures features of global signal patterns. Improved performance in the determination of base methylation can lead to a more accurate diagnosis of an object. Precise measurement of DNA methylation can have several other clinical applications.

[0009] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments relate to systems, devices, and computer-readable media associated with the methods described herein.

[0010] A better understanding of the nature and advantages of the embodiments of the present invention can be obtained by referring to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1Shows an exemplary model framework for DNA methylation according to an embodiment of the present invention.

[0012] Figure 2A And 2B Shows an exemplary measurement window according to an embodiment of the present invention.

[0013] Figure 2C Is a schematic diagram showing an example of integrating the results of a convolutional layer and the position information of DNA fragments to form an input matrix for a downstream layer according to an embodiment of the present invention.

[0014] Figure 2D Illustrates a transformer layer according to an embodiment of the present invention.

[0015] Figure 3A Shows the receiver operating characteristic (ROC) curve comparing models according to an embodiment of the present invention.

[0016] Figure 3B Shows a table comparing the sensitivities of different models according to an embodiment of the present invention at a given specificity.

[0017] Figure 4A Shows the receiver operating characteristic (ROC) curves of different models and datasets according to an embodiment of the present invention.

[0018] Figure 4B Is a graph showing the precision and sub-read depth according to an embodiment of the present invention.

[0019] Figure 5 Is the ROC curve of the HK model 2 using only a single strand according to an embodiment of the present invention.

[0020] Figure 6 Shows the AUC of the methylation analysis of CpG sites at positions relative to the nearest end of sequencing fragments for two datasets according to an embodiment of the present invention.

[0021] Figure 7 Illustrates a scheme for the improved performance of a single-strand model of methylation sites near the 3' end according to an embodiment of the present invention.

[0022] Figure 8A Shows the composition of DNA after TET treatment. According to an embodiment of the present invention, the amounts of 5hmC, 5fC, and 5acC vary with the incubation time.

[0023] Figure 8B Shows the preparation of a 5hmC detection dataset using ligation according to an embodiment of the present invention.

[0024] Figure 8CShows the analytical workflow for the detection of 5mC and 5hmC according to an embodiment of the present invention.

[0025] Figure 9A Shows the ROC curves of the test datasets of the 5xC and 5hmC detectors according to an embodiment of the present invention.

[0026] Figure 9B Shows a block diagram of the modification scores predicted by the 5hmC detector in the test dataset according to an embodiment of the present invention.

[0027] Figure 10A Shows the methylation levels measured in buffy coat and brain samples between different target genomic regions by different methods according to an embodiment of the present invention.

[0028] Figure 10B Shows the methylation levels in human brain samples around the transcription start site (TSS) predicted by the HK model 2 according to an embodiment of the present invention.

[0029] Figure 10C Shows the correlation of 5xC levels in brain samples measured by the HK model 2 and BS-seq according to an embodiment of the present invention.

[0030] Figure 10D Shows the correlation of 5hmC levels (%) in brain samples measured by the HK model 2 and TAB-seq according to an embodiment of the present invention.

[0031] Figure 11A Is a schematic diagram for preparing unmethylated and methylated adenine datasets according to an embodiment of the present invention.

[0032] Figure 11B Shows the IPD distributions in the uA and 6mA datasets according to an embodiment of the present invention.

[0033] Figure 11C Shows the ROC curves of 6mA detection based on the HK model 2 and based only on the IPD metric according to an embodiment of the present invention.

[0034] Figure 11D Shows the false positive rates of 6mA detection based on the HK model 2 and based only on the IPD metric according to an embodiment of the present invention.

[0035] Figure 11E Shows the 6mA methylation levels in non-GATC and GATC environments in Dam-treated DNA samples determined by the HK model 2 according to an embodiment of the present invention.

[0036] Figure 12AShows the IPD distribution in the uC and 4mC datasets according to an embodiment of the present invention.

[0037] Figure 12B Shows the ROC curve of 4mC detection based on the HK model 2 and only based on the IPD metric according to an embodiment of the present invention.

[0038] Figure 13A Shows the 6mA methylation level determined by the HK model 2 according to an embodiment of the present invention.

[0039] Figure 13B Shows the de novo motif analysis related to 6mA modification according to an embodiment of the present invention.

[0040] Figure 14 Illustrates the sparse and dense signal patterns with 6mA modification according to an embodiment of the present invention.

[0041] Figure 15A Illustrates the distribution of kinetic characteristics in the measurement window before and after signal normalization in different datasets according to an embodiment of the present invention.

[0042] Figure 15B Shows the density distribution of kinetic characteristics in different bases of template DNA based on the PacBio Sequel II kit 2.0 according to an embodiment of the present invention.

[0043] Figure 16 Shows the performance of the 6mA classifier using different types of 6mA signals according to an embodiment of the present invention.

[0044] Figure 17 Shows the different sensitivities of different double-strand models at a given specificity according to an embodiment of the present invention.

[0045] Figure 18 Shows the different sensitivities of different single-strand models at a given specificity according to an embodiment of the present invention.

[0046] Figure 19A Shows the HCC methylation fractions in healthy individuals, HBV carriers, and HCC patients determined by the HK model 2 using sequencing DNA molecules with 1 - 6 CpG sites according to an embodiment of the present invention.

[0047] Figure 19B Shows the ROC curve for classifying individuals with or without HCC using the HCC methylation fraction based on molecules with 1 - 6 CpG sites or at least 7 CpG sites according to an embodiment of the present invention.

[0048] Figure 19CShows the pattern of genomic loci relative to the 6mA level in CTCF binding sites according to an embodiment of the present invention.

[0049] Figure 20 Is a flowchart of a process for detecting nucleotide methylation in a nucleic acid molecule according to an embodiment of the present invention.

[0050] Figure 21 Is a flowchart of a process for detecting nucleotide methylation in a nucleic acid molecule according to an embodiment of the present invention.

[0051] Figure 22 Is a graph of detectable CpG sites relative to the distance from the nearest end according to an embodiment of the present invention.

[0052] Figure 23 Shows a workflow for analyzing site methylation using an aptamer sequence according to an embodiment of the present invention.

[0053] Figure 24 Is a graph of the performance of the EMA model for determining the methylation status of CpG sites within 10 nt at the 5' end of a DNA fragment according to an embodiment of the present invention.

[0054] Figure 25 Is a flowchart of a process for detecting nucleotide methylation in a nucleic acid molecule with an aptamer according to an embodiment of the present invention.

[0055] Figure 26 Is a flowchart of a process for detecting nucleotide methylation in a nucleic acid molecule with an aptamer according to an embodiment of the present invention.

[0056] Figure 27 Illustrates a measurement system according to an embodiment of the present invention.

[0057] Figure 28 Is a computer system according to an embodiment of the present invention. Glossary

[0058] "Tissue" corresponds to a group of cells aggregated together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother versus fetus) or to healthy cells versus tumor cells. "Reference tissue" can correspond to the tissue used to determine the tissue-specific methylation level. Multiple samples of the same tissue type from different individuals can be used to determine the tissue-specific methylation level of that tissue type.

[0059] "Biological sample" refers to any sample taken from a subject (e.g., a human (or other animal), such as a pregnant woman, a person with cancer or other disease, or a person suspected of having cancer or other disease, an organ transplant recipient or a subject suspected of having a disease process involving an organ (e.g., the heart in myocardial infarction, or the brain in stroke or the hematopoietic system in anemia)) and containing one or more target nucleic acid molecules. The biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from hydrocele (e.g., testicular), vaginal lavage fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, discharge fluid from the nipple, aspirated fluid from different parts of the body (e.g., thyroid, breast), intraocular fluid (e.g., aqueous humor), etc. Stool samples can also be used. In various embodiments, most of the DNA in a biological sample that has been enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, e.g., greater than 50%, 60%, 70%, 80%, 90%, 95% or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, 3,000 g × 10 minutes, obtaining the liquid portion, and centrifuging again at, for example, 30,000 g for 10 minutes to remove residual cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules of the biological sample can be analyzed (e.g., to provide an accurate measurement). In some embodiments, at least 1,000 cell-free DNA molecules. In other embodiments, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 cell-free DNA molecules or more can be analyzed. At least the same number of sequence reads can be analyzed.

[0060] "Clinically relevant DNA" can refer to DNA of a specific tissue origin to be measured, e.g., to determine the concentration fraction of such DNA or to classify the phenotype of a sample (e.g., plasma). Examples of clinically relevant DNA are fetal DNA in maternal plasma or tumor DNA in patient plasma or other samples with cell-free DNA. Another example includes measuring the amount of graft-related DNA in the plasma, serum or urine of a transplant patient. Other examples include measuring the concentration fraction of hematopoietic and non-hematopoietic DNA in the plasma of a subject, or the concentration fraction of liver DNA fragments (or other tissues) in a sample or the concentration fraction of brain DNA fragments in cerebrospinal fluid.

[0061] "Sequence read" refers to a string of nucleotides obtained from any part or all of a nucleic acid molecule. For example, a sequence read can be a short string of nucleotides (e.g., 20 - 150 nucleotides) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in a variety of ways, such as using sequencing technologies or using probes, such as in hybridization arrays or capture probes that can be used in microarrays, or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification or isothermal amplification using a single primer. Exemplary sequencing technologies include massively parallel sequencing, targeted sequencing, Sanger sequencing, ligation sequencing, ion semiconductor sequencing, and single molecule sequencing (e.g., using nanopores, or single molecule real-time sequencing (e.g., from Pacific Biosciences)). Such sequencing can be random sequencing or targeted sequencing (e.g., by using capture probes that hybridize to specific regions or by amplifying certain regions, both of which enrich such regions). Exemplary PCR techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). As part of the analysis of a biological sample, a statistically significant number of sequence reads can be analyzed. For example, at least 1,000 sequence reads can be analyzed. As other examples, at least 5,000, 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 sequence reads or more can be analyzed.

[0062] "Subread" is a sequence generated from all bases in one strand of a circularized DNA template that has been replicated by a DNA polymerase in a continuous strand. For example, a subread can correspond to one strand of the circularized template DNA. In such an instance, after circularization, a double-stranded DNA molecule will have two subreads: one subread per sequencing run. In some embodiments, the sequence generated can include a subset of all bases in one strand, e.g., due to the presence of sequencing errors.

[0063] "Adaptor" or "adapter" can be an oligonucleotide attached to one end of a nucleic acid molecule. The nucleic acid molecule can be DNA or RNA. Adaptors can include hairpin adaptors, which are oligonucleotides attached to both ends of one end of a double-stranded DNA molecule. Adaptors can facilitate sequencing technologies, including single molecule real-time sequencing. When sequencing a target nucleic acid molecule, some nucleotides of the adaptor can be sequenced.

[0064] A "locus" corresponds to a single locus, which can be a single base position or a set of related base positions, such as a CpG locus. A "locus" can correspond to a region that includes multiple loci. A locus can include only one locus, which would make the locus equivalent to the locus in this context. Various embodiments can analyze a statistically significant number of loci, e.g., at least 100, 200, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more loci.

[0065] "Methylation status" refers to the methylation status at a given locus. For example, a locus can be methylated, unmethylated, or in some cases, indeterminate.

[0066] The "methylation index" of each genomic locus (e.g., a CpG locus) can refer to the proportion of DNA fragments that show methylation at that locus compared to the total number of reads that cover that locus (e.g., as determined from sequence reads or probes). A "read" can correspond to information obtained from a DNA fragment (e.g., the methylation status at a locus). Reads can be obtained using reagents (e.g., primers or probes) that preferentially hybridize to DNA fragments with a specific methylation status at one or more loci. Typically, such reagents are applied after processing using a process that differentially modifies or differentially recognizes DNA molecules based on their methylation status, e.g., bisulfite conversion, or methylation-sensitive restriction enzymes or methyl-binding proteins or anti-methylcytosine antibodies, or single molecule sequencing techniques that recognize methylcytosine and hydroxymethylcytosine (e.g., single molecule real-time sequencing and nanopore sequencing (e.g., from Oxford Nanopore Technologies)).

[0067] The "methylation density" of a region can refer to the number of reads at sites within the region showing methylation divided by the total number of reads covering the sites in that region. The sites can have specific characteristics, such as being CpG sites. Thus, the "CpG methylation density" of a region can refer to the number of reads showing CpG methylation divided by the total number of reads covering CpG sites (e.g., specific CpG sites, CpG sites within a CpG island, or a larger region) in that region. For example, the methylation density per 100-kb bin in the human genome can be determined from the total number of non-converted cytosines (which correspond to methylated cytosines) at CpG sites after bisulfite treatment, as the proportion of all CpG sites covered by sequence reads that map to a 100-kb region. This analysis can also be performed for other bin sizes, such as 500 bp, 5 kb, 10 kb, 50-kb, or 1-Mb, etc. The region can be the entire genome or a chromosome or a part of a chromosome (e.g., a chromosome arm). When the region consists only of that CpG site, the methylation index of the CpG site is the same as the methylation density of the region. The "proportion of methylated cytosines" can refer to the number of cytosine sites "C" that show methylation (e.g., non-converted after bisulfite conversion) in that region compared to the total number of cytosine residues analyzed, i.e., including cytosines outside the CpG context. The methylation index, methylation density, the count of methylated molecules at one or more sites, and the proportion of methylated molecules (e.g., cytosines) at one or more sites are examples of "methylation levels". In addition to bisulfite conversion, other processes known to those skilled in the art can also be used to interrogate the methylation status of DNA molecules, including but not limited to methylation-sensitive enzymes (e.g., methylation-sensitive restriction enzymes), methyl-binding proteins, single-molecule sequencing using platforms sensitive to methylation status (e.g., nanopore sequencing (Schreiber et al. Proc Natl Acad Sci 2013; 110: 18910 - 18915) and by single-molecule real-time sequencing (e.g., those from Pacific Biosciences) (Flusberg et al. Nat Methods 2010; 7: 461 - 465)).

[0068] "Methylomics" provides a measure of the amount of DNA methylation at multiple sites or loci in the genome. Methylomics can correspond to all of the genome, a substantial part of the genome, or a relatively small part of the genome.

[0069] "Pregnancy plasma methylomics" is methylomics determined from the plasma or serum of a pregnant animal (e.g., a human). Pregnancy plasma methylomics is an example of cell-free methylomics because plasma and serum contain cell-free DNA. Pregnancy plasma methylomics is also an example of mixed methylomics because it is a mixture of DNA from different organs or tissues or cells in the body. In one embodiment, such cells are hematopoietic cells, including but not limited to cells of the erythroid lineage (i.e., red blood cells), the myeloid lineage (e.g., neutrophils and their precursors), and the megakaryocyte lineage. During pregnancy, plasma methylomics can contain methylomics information from the fetus and the mother. "Cell methylomics" corresponds to methylomics determined from a patient's cells (e.g., blood cells). The methylomics of blood cells is called blood cell methylomics (or blood methylomics).

[0070] A "methylation signature" includes information related to DNA or RNA methylation at multiple sites or regions. Information related to DNA methylation can include, but is not limited to, the methylation index of CpG sites, the methylation density (abbreviated as MD) of CpG sites in a region, the distribution of CpG sites over a contiguous region, the methylation pattern or level of each individual CpG site within a region containing more than one CpG site, and non-CpG methylation. In one embodiment, a methylation signature can include methylation or non-methylation patterns of more than one type of base (e.g., cytosine or adenine). The methylation signature of a substantial portion of the genome can be considered equivalent to methylomics. "DNA methylation" in the mammalian genome generally refers to the addition of a methyl group to the 5' carbon of a cytosine residue (i.e., 5-methylcytosine) in a CpG dinucleotide. DNA methylation can occur in cytosine in other contexts, such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also be in the form of 5-hydroxymethylcytosine. Non-cytosine methylation, such as N 6 -methyladenine, has also been reported.

[0071] A "methylation pattern" refers to the sequence of methylated and non-methylated bases. For example, a methylation pattern can be the sequence of methylated bases on a single DNA strand, a single double-stranded DNA molecule, or another type of nucleic acid molecule. As an example, three consecutive CpG sites can have any of the following methylation patterns: UUU, MMM, UMM, UMU, UUM, MUM, MUU, or MMU, where "U" represents an unmethylated site and "M" represents a methylated site.

[0072] The terms "hypermethylated" and "hypomethylated" can refer to the methylation density of an individual DNA molecule as measured by the single-molecule methylation level of the individual DNA molecule. For example, the number of methylated bases or nucleotides within the molecule is divided by the total number of methylatable bases or nucleotides within the molecule. A hypermethylated molecule is one in which the single-molecule methylation level is at or above a threshold, which can be defined from application to application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated molecule is one in which the single-molecule methylation level is at or below a threshold, which can be defined from application to application, and which can vary from application to application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%.

[0073] The terms "hypermethylated" and "hypomethylated" can also refer to the methylation level of a population of DNA molecules, as measured by the multi-molecule methylation level of those molecules. A hypermethylated population of molecules is a population in which the multi-molecule methylation level is at or above a threshold, which can be defined from application to application and which can vary from application to application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%. A hypomethylated population of molecules is a population in which the multi-molecule methylation level is at or below a threshold, which can be defined from application to application. The threshold can be 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, and 95%. In one embodiment, the population of molecules can be aligned with one or more selected genomic regions. In one embodiment, the selected genomic region can be disease-related, such as cancer, a genetic disorder, an imprinting disorder, a metabolic disorder, or a neurological disorder. The selected genomic region can have a length of 50 nucleotides (nt), 100 nt, 200 nt, 300 nt, 500 nt, 1000 nt, 2 knt, 5 knt, 10 knt, 20 knt, 30 knt, 40 knt, 50 knt, 60 knt, 70 knt, 80 knt, 90 knt, 100 knt, 200 knt, 300 knt, 400 knt, 500 knt, or 1 Mnt.

[0074] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads that align to that locus. The locus can be as small as a nucleotide, or as large as a chromosomal arm, or as large as an entire genome. The sequencing depth can be expressed as 50x, 100x, etc., where "x" refers to the number of times the locus is covered by sequence reads. The sequencing depth can also be applied to multiple loci or an entire genome, in which case "x" can respectively mean the average number of times the locus or the haploid genome or the entire genome is sequenced. Ultra - deep sequencing can refer to a sequencing depth of at least 100x.

[0075] As used herein, the term "classification" refers to any number or other character associated with a particular property of a sample. For example, the "+" symbol (or the word "positive") can indicate that the sample is classified as having a deletion or an amplification. The classification can be binary (e.g., positive or negative) or have more classification levels (e.g., on a scale from 1 to 10 or 0 to 1).

[0076] The terms "cut - off value" and "threshold" refer to a predetermined number used in an operation. For example, a cut - off size can refer to a size above which fragments are excluded. A threshold can be a value above or below which a particular classification is applied. Either of these terms can be used in any of these contexts. The cut - off value or threshold can be a "reference value" or derived from a reference value representing a particular classification or distinguishing between two or more classifications. The cut - off value can be predetermined with or without reference to the characteristics of the sample or object. For example, the cut - off value can be selected based on the age or sex of the test object. The cut - off value can be selected after the output of the test data and based on the output of the test data. For example, when the sequencing of a sample reaches a certain depth, a certain cut - off value can be used. As another example, a reference object with known classifications and measured characteristic values (e.g., methylation level, statistical size value, or count) of one or more conditions can be used to determine a reference level to distinguish different conditions and / or classifications of conditions (e.g., whether the object has the condition). The reference value can be selected as a representative value between two clusters of a classification (e.g., an average) or a metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, the reference value can be determined based on a statistical simulation of the sample. Either of these terms can be used in any of these contexts. Such reference values can be determined in various ways, as will be understood by a person skilled in the art. For example, metrics can be determined for two different groups of objects with different known classifications, and the reference value can be selected as a representative value between two clusters of a classification (e.g., an average) or a metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, the reference value can be determined based on a statistical simulation of the sample. The specific values of the cut - off value, threshold, reference, etc. can be determined based on the desired precision (e.g., sensitivity and specificity).

[0077] The term "cancer grade" can refer to the presence or absence of cancer (i.e., present or absent), the stage of cancer, the size of a tumor, the presence of metastasis, the total tumor burden of the body, the response of the cancer to treatment, and / or other measures of the severity of the cancer (e.g., recurrence of the cancer). The cancer grade can be a number or other marker, such as a symbol, letter, and color. The grade can be zero. The cancer grade can also include pre-deterioration or pre-cancerous conditions (states). The cancer grade can be used in various ways. For example, screening can check for the presence of cancer in people who were previously not known to have cancer. An assessment can investigate people who have been diagnosed with cancer to monitor the progression of the cancer over time, study the effectiveness of treatment, or determine a prognosis. In one embodiment, the prognosis can be expressed as the chance that a patient will die of cancer, or the chance or degree that the cancer will progress after a specific duration or time, or the chance or degree of cancer metastasis. Detection can mean'screening' or can mean checking whether someone with suggestive features of cancer (e.g., symptoms or other positive tests) has cancer.

[0078] "Pathology grade" (or disease grade) can refer to the amount, degree, or severity of pathology associated with an organism, where the grade can be as described above for cancer. Another example of pathology is the rejection of a transplanted organ. Other exemplary pathologies can include genomic imprinting disorders, autoimmune attacks (e.g., lupus nephritis or multiple sclerosis that damage the kidneys), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). The health state of an object can be considered a classification of no pathology.

[0079] "Pregnancy-related disorders" include any disorder characterized by an abnormal relative expression level of genes in maternal and / or fetal tissues. These disorders include, but are not limited to, preeclampsia, intrauterine growth restriction, invasive placentation, preterm birth, hemolytic disease of the newborn, placental insufficiency, fetal hydrops, fetal malformations, HELLP syndrome, systemic lupus erythematosus, and other immune diseases of the mother.

[0080] The abbreviation "bp" refers to base pairs. In some cases, "bp" can be used to represent the length of a DNA fragment, although the DNA fragment can be single-stranded and does not include base pairs. In the context of single-stranded DNA, "bp" can be interpreted as providing the length in nucleotide units.

[0081] The abbreviation "nt" refers to nucleotides. In some cases, "nt" can be used to represent the length of single-stranded DNA in base units. Also, "nt" can be used to represent a relative position, such as upstream or downstream of a locus being analyzed. In some environments involving technical conceptualization, data representation, processing, and analysis, "nt" and "bp" can be used interchangeably.

[0082] The term "sequence context" can refer to the base composition (A, C, G, or T) and base order in a stretch of DNA. This stretch of DNA can surround the base for which base methylation analysis is to be performed or serve as the target for base methylation analysis. For example, the sequence context can refer to the bases upstream and / or downstream of the base for which base methylation analysis is to be performed.

[0083] The term "kinetic feature" or "kinetic signature" can refer to features obtained from sequencing, including from single molecule real-time sequencing. Such features can be used for base methylation analysis. Exemplary kinetic features include upstream and downstream sequence context, strand information, interpulse duration, pulse width, and pulse intensity. In single molecule real-time sequencing, the effect of polymerase on a DNA template is continuously monitored. Thus, the measurements resulting from such sequencing can be considered kinetic features, such as nucleotide sequences.

[0084] A “machine learning model” (ML model) can refer to a software module configured to run on one or more processors to provide classification or numerical characteristics of one or more samples. Sample data (e.g., training data) can be used to generate an ML model to make predictions on test data. An example is an unsupervised learning model. Another example type of model is supervised learning that can be used with embodiments of the present disclosure. Exemplary supervised learning models can include different methods and algorithms, including analytical learning, statistical models, artificial neural networks, backpropagation, boosting (meta-algorithm), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum description length (decision trees, decision diagrams, etc.), multilinear subspace learning, naive Bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probably approximately correct (PAC) learning, link wave descent rules, knowledge acquisition methodologies, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, minimal complexity machines (MCMs), random forests, classifier ensembles, ordinal classification, data preprocessing, handling imbalanced datasets, statistical relational learning, or Proaftn — a multi-criteria classification algorithm. Models can include linear regression, logistic regression, deep recurrent neural networks (e.g., long short-term memory, LSTM), hidden Markov models (HMMs), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering of applications with noise (DBSCAN), random forest algorithms, support vector machines (SVMs), or any of the models described herein. Various cost / loss functions that define the error from known labels (e.g., sum of squared and absolute differences from known classifications) and various optimization techniques, such as using backpropagation, steepest descent, conjugate gradient, and Newton and quasi-Newton techniques, can be used to train supervised learning models in various ways.

[0085] The term “deep learning” can refer to the use of artificial neural networks with multiple layers in a network. The number of layers can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40 or more, or any number within the ranges and including these numbers between these numbers.

[0086] The term “transformer” or “transformer layer” can refer to a machine learning model that employs a self-attention mechanism to differentially weight the importance of each part of the input data. Transformers are designed to process sequential input data simultaneously rather than sequentially. Transformers are described in Vaswani et al., “Attention is All You Need,” arXiv:1706.03762 (2017).

[0087] The term "real-time sequencing" can refer to techniques involving data collection or monitoring during the course of a reaction involved in sequencing. For example, real-time sequencing can involve optical monitoring or imaging of a DNA polymerase incorporating new bases. As another example, real-time sequencing can involve monitoring of an electrical signal of an ionic current through a nanopore as a nucleotide strand translocates through the nanopore.

[0088] The term "electrical signal" can refer to a voltage or current that conveys information. Electrical signals can be represented in various regular and / or irregular signal waveform types and / or shapes, such as square waves, rectangular waves, triangular waves, sawtooth waveforms, or various pulses and spikes. Electrical signals can include a visual representation of a voltage or current varying over time. Measurements of an electrical signal can be sampled at a particular time (e.g., milliseconds). For example, current is sampled at frequencies of 1 kHz, 2 kHz, 3 kHz, 4 kHz, 5 kHz, 10 kHz, 20 kHz, 30 kHz, 40 kHz, 50 kHz, 100 kHz, etc.

[0089] The term "signal segment" or "segment" can refer to a portion of a trace of an electrical signal associated with a particular nucleotide sequencing. The segment can correspond to a nucleotide determined from base detection in nanopore sequencing. The segment can cover a certain duration of the trace. Different segments can have different durations. The segments can be non-overlapping. In some embodiments, the electrical signal amplitude can have some variation within the segment. For example, the electrical signal amplitude can be within 5%, 10%, 20%, 30%, or 40% of the average or median electrical signal amplitude within the segment.

[0090] The term "about" or "approximately" can mean within an acceptable error range of a particular value determined by a person of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations according to the practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within one order of magnitude, within 5-fold, and more preferably within 2-fold of a value. When a particular value is described in this application and the claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable error range of that particular value. The term "about" can have the meaning commonly understood by a person of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.

[0091] Standard abbreviations can be used, e.g., bp, base pair; kb, kilobase; pl, picoliter; s or sec, second; min, minute; h or hr, hour; aa, amino acid; nt, nucleotide, etc.

[0092] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments of this disclosure, some potential and exemplary methods and materials are now described. Detailed Description of the Invention

[0093] There is a need for accurate and effective methods for detecting base methylation without enzymatic or chemical conversion (e.g., bisulfite conversion). Other methods and systems for directly detecting base methylation can depend on specific technologies (e.g., certain reagents or protocols) or equipment. The embodiments described herein use machine learning models to detect base methylation, which can be applied to a wide range of technologies or equipment. Embodiments can include machine learning models that capture local patterns (e.g., convolutional layers) and global patterns (e.g., transformer layers). Local patterns can be patterns generated by elements in a given convolutional filter. Global patterns can be patterns generated outside a given convolutional filter, and the distances of these global patterns can be up to the size of the measurement window. Additionally, previous methods for detecting base methylation may not be suitable for detecting methylation at or near the ends of DNA fragments. The methods described herein can use signals from adaptors at the ends of DNA fragments to assist in detecting methylation.

[0094] The embodiments presented in this disclosure can be used for DNA obtained from, but not limited to, the following: cell lines, samples from organisms (e.g., solid organs, solid tissues from pregnant women, samples obtained via endoscopy, blood or plasma or serum or urine, chorionic villus biopsy, etc.), samples obtained from the environment (e.g., bacteria, cell contaminants), food (e.g., meat). In some embodiments, the methods presented in this disclosure can also be applied after a step in which a portion of the genome is first enriched, such as using hybridization probes (Albert et al., 2007; Okou et al., 2007; Lee et al., 2011), or methods based on physical separation (e.g., based on size, etc.) or subsequent restriction enzyme digestion (e.g., MspI) or Cas9-based enrichment (Watson et al., 2019). Although the present invention does not require enzymatic or chemical conversion to work, in certain embodiments, such conversion steps can be included to further enhance the performance of the present invention.

[0095] Embodiments of the present disclosure allow for improved accuracy, utility, or convenience in detecting base methylation or measuring methylation levels. Methylation can be detected directly. Embodiments can avoid enzymatic or chemical conversions that may not retain all methylation information for detection. Additionally, certain enzymatic or chemical conversions may be incompatible with certain types of methylation. Embodiments of the present disclosure can also avoid amplification by PCR, which may not transfer base methylation information to the PCR product. Additionally, both strands of DNA can be sequenced together, enabling pairing of the sequence from one strand with its complementary sequence on the other strand. In contrast, PCR amplification separates the two strands of double-stranded DNA, making such sequence pairing difficult.

[0096] Methylation signatures determined with or without enzymatic or chemical conversion can be used to analyze biological samples. In one embodiment, methylation signatures can be used to detect the origin of cellular DNA (e.g., maternal or fetal, tissue, viral, or tumor). Detection of abnormal methylation signatures in tissues aids in the identification of developmental disorders and the identification and prediction of tumors or malignancies in an individual. Imbalances in methylation levels between haplotypes can be used to detect disorders, including cancer. Methylation patterns in individual molecules can identify chimeras (e.g., between virus and human) and hybrid DNA (e.g., between two genes that are not normally fused in the native genome); or between two species (e.g., through genetic or genomic manipulation).

[0097] Methylation analysis can be improved through enhanced training, which can include narrowing the data used in the training set. Specific regions can be targeted for analysis. In embodiments, such targeting can involve enzymes that can cleave DNA sequences or genomes based on DNA sequences or genomic sequences, either alone or in combination with other reagents. In some embodiments, the enzyme is a restriction enzyme that recognizes and cleaves specific DNA sequences. In other embodiments, more than one restriction enzyme with different recognition sequences can be used in combination. In some embodiments, the restriction enzyme can cleave or not cleave based on the methylation status of the recognition sequence. In some embodiments, the enzyme is an enzyme within the CRISPR / Cas family. For example, the CRISPR / Cas9 system or other systems based on guide RNAs (i.e., short RNA sequences that bind to complementary target DNA sequences and in the process direct the enzyme to act on the target genomic location) can be used to target target genomic regions. In some cases, methylation analysis is possible without alignment to a reference genome. I. Machine Learning Techniques for Detecting Methylation

[0098] The embodiments described herein include a transformer layer combined with a convolutional layer to detect base methylation. The transformer layer uses an attention mechanism to encode data resulting from one or more convolutional layers. A transformer is a deep learning model that utilizes a self-attention mechanism to differentially weight the importance or salience of each part of the input data in the current context. Similar to cognitive attention, attention can enhance some parts of the input data while reducing others, enabling the classification model to focus more on small but important parts of the data. Learning which parts of the data are more important than others depends on the context and can be trained using gradient descent in the backpropagation process. The self-attention mechanism allows the inputs to interact with each other ("self") and identify the parts of the data to which the model should pay more attention ("attention"). Thus, a model with attention can allow for fast learning and can achieve more accurate predictions by enhancing the influence of those more relevant input parts while reducing the influence of those less relevant input parts.

[0099] An important process in the transformer is to transform the original input data matrix into three data matrices (i.e., the query matrix (Q), the key matrix (K), and the value matrix (V)) by multiplying with the corresponding weight matrices W Q , W k and W v . The product between Q and K can be calculated to form the attention scores (also known as the attention filter S0) for each part of the input. Then, the attention filter (S0) is multiplied with the value matrix (V), i.e., S0·V, to obtain the self-attention scores (S), which assign high weights to the features that are more important for classification accuracy. One run of the above process is called single-head attention. Multiple runs of this process can be applied to obtain different groups of relevant features to which the model should pay different attention, called multi-head attention. We provide data showing that a machine learning model with a transformer layer is more accurate than a model without a transformer layer. A. Model Framework

[0100] Figure 1 Shows an exemplary model framework for DNA methylation using the kinetic signals of DNA polymerase during single-molecule sequencing. Single-molecule sequencing can include single-molecule real-time (SMRT) sequencing or nanopore sequencing.

[0101] In stage 104, the values of the kinetic signals ( Figure 1 "kinetic values" in Figure 1What is shown as "base information" in [figure] is organized into an input layer, which can be a digital matrix or a vector. "Position information" refers to the relative position of a base on a strand. These kinetic signals can be any of the kinetic features described herein and are described in more detail elsewhere in the present disclosure, including Sections I.A.1 and II.A. Although Stage 104 shows the use of Watson strand data and Crick strand data, a single strand can be used instead of two strands.

[0102] In Stage 108, the input layer is processed by one or more locally connected layers (e.g., a convolutional layer that shares weights among a set of local sequence base positions). As an example, two one-dimensional (1D) convolutional layers can be used. As another example, 128 filters with a kernel size that includes all kinetic signals from 5 positions (i.e., 5 nucleotide positions in the measurement window) are applied to each convolutional layer (also referred to as the latent dimension). Each position contributes two types of kinetic signals, namely IPD and PW. When using a 1-D convolutional layer, all kinetic signals across rows from 5 positions are automatically included. Thus, in this example, the kernel size can be equal to 8×10. In some embodiments, the kernel size can include, but is not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or more nucleotide positions, or any combination thereof. In some other embodiments, the number of filters can be, but is not limited to, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, 500, or a combination thereof. The number of filters can be within the range between any two numbers and includes any two numbers.

[0103] A batch normalization layer can be applied between convolutional layers using a rectified linear unit (Relu) activation function. Batch normalization can standardize the input of a layer across a mini-batch by recentering and rescaling. A mini-batch refers to a subset of training samples. For example, batch normalization will process the values (represented by x) in a matrix based on the overall mean (represented by m) and standard deviation (represented by sd) according to the formula (x - m) / sd. Thus, the mean and standard deviation of the input are close to 0 and 1, respectively. In the example, the output of one convolutional layer can be the input to a second convolutional layer. Each resulting output from the convolutional layer can be integrated with the position information in the DNA fragment being analyzed.

[0104] In stage 112, the convolution result is added to the location information (location embedding 120) as the input to the transformer layer. The transformer layer uses the self-attention mechanism to generate the transformer result that takes into account the global signal pattern.

[0105] In stage 116, the transformer result is processed by the output layer. The output layer processing can generate the probabilities of methylation (circle 124) and unmethylation (circle 128). Activation functions can be used to provide the probabilities. Examples of activation functions include softmax activation function, ridge activation function (e.g., linear activation, ReLU activation, heaviside activation, logistic activation), radial activation function (e.g., Gaussian, multiquadric, inverse multiquadric, multiharmonic spline), sigmoid, identity, and binary step. A methylation probability greater than the threshold indicates methylation.

[0106] The output layer in stage 116 includes new data types for determining methylation. Figure 1 A method for generating these new data types is shown. The method can include specific machines for Figure 1 one or more stages. These specific machines can include a computer with a specific processor. For example, such a processor can include a central processing unit (CPU) and a graphics processing unit (GPU) designed to accelerate computer graphics and image processing. In some embodiments, the computer can include field-programmable gate arrays (FPGAs) configured for machine learning detection of methylation.

[0107] In some embodiments, the probability of methylation can include the probability of a specific type of methylation. For example, the probability of methylation can include the probability of 5mC methylation, the probability of 5hmC methylation, the probability of 6mA methylation, or the probability of any other type of methylation. For multiple types of methylation, multiple circles can replace circle 124. 1. Input layer

[0108] In some embodiments, the activation function may include, but is not limited to, binary step function, linear activation function, non-linear activation function, sigmoid / logistic activation function, hyperbolic tangent, exponential linear unit (ELUs) function, Swish function, or Gaussian error linear unit (GELUs). The output layer may include a plurality of neurons, where multiple arithmetic operations are performed, such as multiplications by the weights and additions by biases. The number of neurons may be, but is not limited to, 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1000. The number of neurons may be in the range between any two numbers and include any two numbers. The number of output layers may be, but is not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, etc. The number of output layers may be in the range between any two numbers and include the number of any two numbers. In some embodiments, 2-D, 3-D convolutional layers or other combinations may be used, but are not limited to them.

[0109] The kinetic signals of DNA polymerase may include inter-pulse duration (IPD) and pulse width (PW). IPD is a measure of the length of the time period between two emission pulses, each emission pulse indicating the incorporation of a different fluorescently labeled nucleotide into the nascent strand. PW is another measure that reflects polymerase kinetics and is related to the pulse duration associated with base incorporation. The kinetic signals corresponding to a series of sequenced nucleotides can be organized into a matrix (referred to as a measurement window) based on the relative sequence position.

[0110] Figure 2A An exemplary measurement window showing the combination of Watson strand data and Crick strand data in a single two-dimensional (2-D) data matrix is shown. The measurement window may include kinetic signals, including PW and IPD values from 10-nt upstream and 10-nt downstream of the cytosine targeted for methylation analysis. The first column of the matrix represents the type of nucleotide being studied. In the first row of the matrix, the position of 0 represents the target base for base methylation analysis. The relative positions of -1, -2, and -3 indicate positions 1-nt, 2-nt, and 3-nt upstream of the base for which base methylation analysis is performed, respectively. The relative positions of +1, +2, and +3 indicate positions 1-nt, 2-nt, and 3-nt downstream of the base for which base methylation analysis is performed, respectively. Each position includes two columns, which contain the corresponding IPD and PW values.

[0111] The four rows after the row with the IPD and PW headers correspond to the four types of nucleotides (A, C, G, and T) in a strand (e.g., the Watson strand). The presence of IPD and PW values in the matrix depends on which corresponding nucleotide type is sequenced at a particular position. For example, as Figure 2A shown, at relative position 0, the IPD and PW values are shown in the row representing "G" in the Watson strand, indicating that guanine is detected in the sequence result at that position. Other grids in columns not corresponding to the sequenced bases can be encoded as "0". As an example, as Figure 2A shown, the sequence information corresponding to the two-dimensional (2-D) data matrix ( Figure 2A ) is 5'-GATGACT-3' of the Watson strand. A similar process for constructing the measurement window can be applied to the data generated from the Crick strand.

[0112] The measurement window can be used as the input layer for initializing model training and testing. In some embodiments, the length of the DNA upstream of the locus for base methylation analysis can be, but is not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or 1000 nucleotides. In some embodiments, the length of the DNA downstream of the locus for base methylation analysis can be, but is not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, etc. The lengths upstream and downstream of the locus for base methylation analysis can be equal or unequal. The length can be in the range between any two of the disclosed numbers and includes any two of the disclosed numbers.

[0113] In some embodiments, only one of the Watson strand or the Crick strand can be used. A model using single-strand data as input can detect methylation. Using only a single strand is discussed in more detail in Section I.D.

[0114] Figure 2BShows an example of a measurement window of data from a single strand with base encoding. As Figure 2B shown, base information can be encoded by different numbers: for example, A = 0, T = 1, C = 2, and G = 3. Such numerical base information can be organized as shown at the bottom of a 2-D data matrix. Data from the Watson strand and the Crick strand can be used as two independent inputs. The data can then be processed with convolutional layers and transformer layers, as Figure 1 shown. In an embodiment, the processed information matrices obtained from the Watson and Crick strands can be vertically (e.g., as done for Figure 2A ) or horizontally (as Figure 2B shown) concatenated and further input into an output layer to generate a methylation probability.

[0115] The size of the measurement window can be, but is not limited to, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, 35 nt, 36 nt, 37 nt, 38 nt, 39 nt, 40 nt, 50 nt, 60 nt, 70 nt, 80 nt, 90 nt, 100 nt, 200 nt, 500 nt, or any range between any two numbers and including any two numbers.

[0116] Some sequencing techniques (e.g., PacBio SMRT sequencing) involve ligating an adapter sequence to a DNA fragment. Typically, the information associated with the adapter sequence is trimmed so that the resulting sequence reads for the end user do not have the adapter. Thus, many CpG sites near the 5' end or 3' end of a DNA sequence may not have sufficient flanking kinetic signals for methylation analysis. For example, for a measurement window size of 21 centered on a cytosine, CpG sites within 10 nt of the 5' end will not have sufficient flanking kinetic signals for methylation analysis. In an embodiment, the kinetic signals from the trimmed adapter sequence can be used and incorporated into the kinetic signals from the target DNA fragment so that any CpG site in the sequenced DNA fragment can be analyzed, even near the end of the fragment.

[0117] The measurement window is described in U.S. Patent Publication No. 2021 / 0047679A1, filed Aug. 17, 2020, the entire content of which is incorporated herein by reference for all purposes. 2. Dropout Layer

[0118] In some embodiments, a dropout layer can be applied between any of the layers mentioned in each stage to improve the generality of the model and prevent overfitting. The dropout layer allows some features or neurons to be ignored before being input to the next layer. The dropout layer can be implemented in such a way that during training, at each step, the input units of each layer are randomly set to 0 by a certain dropout rate (e.g., if a 20% dropout rate is used, 20% of the neurons will be ignored and marked as 0). 3. Convolutional Layer

[0119] The convolutional layer contains a set of filters (or kernels) whose parameters are learned from training. The size of the filter is usually smaller than the actual size of the measurement window. Each filter slides horizontally and vertically between the input matrices, and at each spatial position, the dot product between the filter and the input is calculated, generating a convolutional output. For example, if the filter is 3×3, the output can be the sum or average of the input values, and the output is assigned to an intermediate node. The convolution process can be repeated with different filters. Such a process can capture local patterns of adjacent signals (i.e., local signal patterns).

[0120] Figure 2C An exemplary process of integrating the results of the convolutional layer and the position information of the DNA segments to generate an input matrix for the downstream transformer layer is shown. As Figure 2C shown, the original input matrix 250 can be a matrix with a dimension (i.e., shape) of 8×42 (e.g., Figure 2A ). Matrix 250 includes 8 rows indicating A, C, G, and T from each of the Watson and Crick strands and 42 columns indicating the IP and PW values between positions in the 21-nt measurement window. For the clarity of the whole figure, matrix 250 does not show all rows and columns to scale.

[0121] For example, a filter kernel having a size of 8×6 (i.e., all bases and 3 positions, including IP and PW at each position) can scan a measurement window from left to right with a sliding step size (e.g., 1). Processing by the convolutional layer can generate an output matrix of size 1×42 with padding outside the edges (e.g., an 8×5 zero-padding matrix). This padding operation can allow the size of the columns to be the same in both the input matrix and the output matrix. In some other embodiments, the filter kernel size can be n×m, where n can be, but is not limited to, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, or any number within the range between these numbers and including these numbers, and m can be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, or any number within the range between these numbers and including these numbers. The sliding step size can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or any number within the range between these numbers and including these numbers. The filter kernel can have a filter kernel size smaller than the size of the window used to create the data structure.

[0122] Different filters can contain a different set of weights, which can produce different convolutional layers 254. One or more filters (e.g., having the same size) can be used, and one or more output matrices can be generated in the convolutional layer. The number of filters can be, but is not limited to, 1, 2, 3, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or any number within the range between these numbers and including these numbers. As an example, there can be 128 convolutional filters.

[0123] Based on the position information in the measurement window, called the convolutional output, all outputs from the convolutional layer are concatenated and organized in a matrix 258 with a latent dimension. The latent dimension depends on the size of the input matrix, the number of filters, and the size of each filter used (e.g., if no padding is performed). The convolutional results (e.g., the convolutional output matrix 258) generated by the sliding filter kernel can be indexed by the nucleotide positions in the measurement window. The position indices can be stored in another matrix 262 representing the relative position relationships. By embedding, the matrix 258 containing the convolutional output and the matrix 262 containing the position indices are processed. For example, if there are 128 convolutional filters, the matrix 258 and the matrix 262 can have dimensions 128×42.

[0124] The position index in matrix 262 can specify the position in the DNA segment relative to the CpG site under discussion. In this example, the relative position ranges from 10 nt upstream to 10 nt downstream of the CpG site.

[0125] Embedding refers to digital operations regarding spatial transformation. For example, mapping an n-dimensional vector (or matrix) to an m-dimensional vector (or matrix), including linear transformation, rotation, and scaling. Compared with the data matrix before embedding, embedding can increase or decrease the dimension. Compared with traditional dimensionality reduction techniques such as principal component analysis (PCA), an important feature of the embedding in this article is that the embedding space (e.g., the weights used in linear transformation) can be learned according to backpropagation and the loss function. As a result, adding an embedding layer can improve the model performance.

[0126] As Figure 2C shown, the convolutional signal embedding 266 can analyze and store the data from the convolutional layer. 4. Position Index

[0127] The position embedding 270 can store the position index in a matrix with the same shape (dimension) as that used in the convolutional signal embedding 266. The relative position information can be encoded through periodic functions (e.g., sine, cosine, tangent, cotangent) as the initial weights, and the relative position information can be learned during the training of the model.

[0128] The position embedding can incorporate information related to the time series during sequencing to consider the order of features. The position embedding helps the model identify the pattern information at any position in the feature map, which can be obtained from the CNN process. The position embedding can help the model in capturing the pattern relationships between any positions in the feature map through the self-attention mechanism. The position embedding provides a space containing the feature map such that the positions in the feature map are trainable, enabling effective learning of the pattern information between positions.

[0129] The input matrix (I) 274 for input into the transformer layer will include the results from the convolutional signal embedding and the position embedding. In the original input matrix 250, the states in the columns of the measurement window are vectors corresponding to the base position i. The input matrix (I) 274 is after the convolutional and embedding layers. In the input matrix (I) 274, column X can correspond to the base position i. X is a vector with the convolutional output as elements. The input matrix (I) can have the same dimension as matrix 258 and matrix 262 (e.g., 128×42).

[0130] The position encoding process through the sine / cosine function can refer to the following formula for calculating the position information in the latent dimension: And Where (p, 2i) represents the position index p on the 2i-th axis (even axis), and (p, 2i + 1) represents the position index p on the (2i + 1)-th axis (odd axis); l represents the size of the potential dimension (e.g., the number of rows) of the cascaded convolution result. 5. Transformer Layer

[0131] One or more Transformer layers can be applied to the intermediate result, for example, after applying a locally connected layer, such as one or more convolutional layers. As an example, two or three Transformer layers can be applied after a Convolutional Neural Network (CNN). The Transformer layer (also referred to as a block in some contexts) can use a self-attention mechanism that allows differential weighting of the relevance of each part of the input data at all positions. Each Transformer block can have a series of parameters for training, including the number of multi-head self-attention, QKV biases (queries, keys, and values in the attention mechanism), and the ratio of the multi-layer perceptron. Multi-head self-attention refers to a process by which the attention mechanism can be applied several times in parallel on the data input matrix. In one example, the number of multi-head self-attention is 4. In some embodiments, the number of multi-head self-attention can be 1, 2, 3, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or 50, or any range between any two numbers and including any two numbers.

[0132] In the Transformer model, multiple numerical operations are used to implement the attention mechanism. The self-attention module takes n inputs and returns n outputs, allowing the inputs to interact with each other (i.e., "self") to determine which inputs should be given higher weights (i.e., "attention"). The output is a set of these interactions and attention scores. In one embodiment, the numerical operations can be vectorized. In some cases, vectors and matrices can be converted to each other. For example, a matrix can be composed of multiple vectors. As another example, a matrix can be indexed into different vectors.

[0133] As an example, the input vector of the Transformer layer in the input matrix (I) 274 can be a layer, for example, the intermediate result after the operation of the convolutional layer 254 as shown. As another example, a series of input vectors can correspond to the columns in the measurement window, for example, as shown, without applying a convolutional layer. Each column corresponds to the state of the nucleic acid at that position. Figure 2C As another example, a series of input vectors can correspond to the columns in the measurement window, for example, as shown, without applying a convolutional layer. Each column corresponds to the state of the nucleic acid at that position. Figure 2A shown, without applying a convolutional layer. Each column corresponds to the state of the nucleic acid at that position.

[0134] Figure 2D Shows how the input matrix (I) 274 is processed by the Transformer layer. The input matrix (I) has dimensions, for example, 42×128. Note that the matrix dimensions are swapped from Figure 2C swapped. Figure 2D The matrix 274 in can be Figure 2CThe transpose of matrix 274 in. The input vectors in input matrix (I) 274 can interact with the weighted matrices (W Q , W K and W V ) 278A, 278B, 278C corresponding to three representations including the secret key (K) matrix, the query (Q) matrix, and the value (V) matrix. For each input vector, there are W Q , W K and W V , and they can be regarded as parameters, such as the weights of a neural network. Each weighted vector can be multiplied by the corresponding input vector to provide the representation vectors K, Q, and V (280B, 280A, 280C).

[0135] Regarding the values of these parameter matrices (K, Q, and V), the weights (values) used in W Q , W K and W V can be randomly initialized. Each input matrix can be multiplied by a set of weighted matrices W Q , W K and W V respectively, and each intermediate matrix result can be added to the corresponding bias matrix (B) to generate K, Q, and V. The K 280B, Q 280A, and V 280C matrices can have the same dimension as the input matrix (I), such as 42×128.

[0136] For the input matrix corresponding to base position I, Q i can be used as a query for the secret key matrix of all input matrices (including K i and another input matrix) at each position. For example, the inner product can be applied between Q i and the corresponding K matrix. The obtained value can be normalized on a scale between 0 and 1 (for example, using an activation function). Then, the normalized value can be multiplied by the value V matrix and summed to obtain the self-attention score matrix.

[0137] This multiplication can be performed many times, such as but not limited to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400, or 500 times, or any range between any two numbers and including any two numbers. In an embodiment, for a given input matrix (I) 274, the transformer layer can be implemented as follows: 1. Prepare the input (represented by I) generated based on the relevant detailed description of Figure 2C , which is derived from the convolutional layer and combined with the position information. The convolutional layer processes the original measurement window with dynamic signals including IPD and PW metrics, as well as the identity of the base. 2. Standardize the I application layer and then linearly transform it into three matrices, respectively called the query matrix (Q), the key matrix (K), and the value matrix (V). For example, "Q = I · W Q + B Q " can be set, where "·" represents the digital operation of dot product; "W Q " represents the weight matrix; and "B Q " represents the digital bias added to this linear transformation. Similarly, K = I · W K + B K and V = I · W V + B V can be set. Layer normalization normalizes each input in the batch across all features, which is different from batch normalization that normalizes each feature across the batch. 3. Calculate the running attention score S0 (phase 284) of a process or self-attention layer according to the following exemplary formula: where (Q × K T ) represents the matrix multiplication between the Q and K matrices in this practice. In another embodiment, cosine similarity can be used to calculate the attention score. Any similarity measure or projection between two vectors can be used. D K represents the dimension of the matrix K (which is equal to the dimensions of Q and V), and dividing by the square root of d K represents the scaling factor for additional attention. The attention score S0 is a matrix and is discussed in more detail below. 4. Calculate a self-attention score (S) (phase 286) by multiplying the attention score by the value matrix (V). One head of self-attention refers to one run of the self-attention layer. Multi-head attention will produce several different self-attention scores (i.e., S 0 , S 1 , … S n ). In this case, all self-attention scores can be concatenated to obtain a merged matrix (phase 288) as the intermediate output of the self-attention process. The self-attention score is described in more detail below. 5. A set of fully connected neural network layers (phase 290) (possibly along with multi-layer perceptron [MLP] layers) to obtain the output (O) 292 of the transformer layer. 6. When the above steps are repeated under this set of transformer layers, an output can be obtained that can be used as the downstream input of a neural network layer (e.g., fully connected) for generating methylation probability.

[0138] The attention score S0 is a matrix with a dimension of n×m, rather than a single value. In one instance, n can be the measurement window size × 2. If a 21-nt measurement window is used, each position has two kinetic features, namely IPD and PW. In Figure 2D 's instance, n can be 42 (21×2). M can depend on the number of filter kernels, the filter kernel size, and padding. For example, if 128 filter kernels with a size of 8×6 covering 3 nucleotide positions are used in each step of convolution and padding is allowed, the output can be obtained during convolution without changing the column dimension [e.g., the dimension of the output in convolutional layer 254 is (1×42)]. The convolution results can be concatenated so that a 42×128 input matrix (I) for the transformer layer can be obtained. When embedding processing is applied to the input matrix (I) 274, the dimension of the input matrix 274 can remain unchanged. As Figure 2D shown, matrices W Q , W K , and W V can be used to apply matrix transformation to the input matrix (I) 274 to obtain Q, K, and V (280A, 280B, 280C) matrices, each with a dimension of 42×128. The softmax function is a function that converts a vector of K real values into a vector of real values with a sum of 1. The formula is shown below: where V represents a vector containing k elements, e.g., V = (V0, V1, …, V k ). Therefore, the attention score (S0) obtained from the matrix multiplication of Q and K T is a matrix with a dimension of 42×42.

[0139] The self-attention score S 0 can be obtained by S0·V and has a dimension of 42×128. If 4 multi-head self-attention operations are performed, the concatenation of the multi-head self-attention results can lead to a 42×512 matrix, which can be further applied through a fully connected layer.

[0140] In some embodiments, a pooling operation, such as but not limited to max pooling, average pooling, etc., can be performed on the convolution results to replace concatenation or supplement concatenation. In Figure 2D , if pooling is applied, the Q, K, and V matrices will each be 42×1. At stage 286, S 0 also has a dimension of 42×1. Using 4 multi-head self-attention, stage 288 will produce a 42×4 combined matrix.

[0141] Layer normalization can be applied before and after each self-attention layer. In stage 290, the MLP layer can include a GELU layer to improve performance. GELU represents the Gaussian Error Linear Unit - a type of activation function. GELU can weight the input according to the value of the input, rather than gating the input as in the Rectified Linear Unit (ReLU). In one example, GELU can be defined as follows: where 'erf' is the error function defined as follows

[0142] The methods described herein integrate convolutional networks and transformer architectures. These methods can facilitate the modeling of local and global signal patterns, resulting in improved methylation analysis. 6. Output Layer

[0143] The softmax activation function can be utilized, and an output layer with four neurons can be applied to produce the probability scores of methylated CpG sites (i.e., the probability of methylation), resulting in the output of the probability of methylation 124 and the probability of non-methylation (circle 128). In one embodiment, the sigmoid activation function can be used to produce the probability scores of methylated CpG sites (i.e., the probability of methylation), resulting in the output of the probability of methylation (circle 124) and the probability of non-methylation (circle 128). In some embodiments, the probability of methylation (circle 124) can include several different probabilities, each for a different type of methylation. For example, the probability of 5mC methylation can be estimated. The probability of 5hmC methylation, 6mA methylation, or any other type of methylation can be estimated. B. Training and Test Datasets

[0144] To demonstrate the effectiveness of the proposed model using convolutional neural networks and transformer layers for methylation analysis (referred to herein as the enhanced methylation analysis [EMA] model or also as HK model 2), we generated training datasets, including an unmethylated dataset and a methylated dataset. The unmethylated dataset contained sequencing results from amplified DNA prepared via whole genome amplification (WGA) (referred to as the WGA dataset). The use of unmodified nucleotides in WGA resulted in amplified DNA containing little base methylation (except for a small amount of input genomic DNA). The methylated dataset contained sequencing results from DNA treated with M.SssI prior to sequencing (referred to as the M.SssI-treated dataset). M.SssI is a CpG methyltransferase isolated from an Escherichia coli strain containing a methyltransferase gene from the Sprioplasma sp. strain MQ1. The M.SssI methyltransferase methylates CpG sites in double-stranded DNA (Greer et al., Cell. 2015; 161:868-878). In the dataset of the samples treated with M.SssI, half of the sequenced CpG sites were used to train the overall kinetics (HK) model (Tse et al. Proc Natl Acad Sci USA. February 2, 2021; 118:e2019768118). Among other differences, the HK model does not include the transformer layer of the EMA model. In the WGA dataset, an equal number of CpG sites were randomly sampled to train the HK model. The remaining half of the CpG sites in the dataset of the M.SssI-treated samples and the same number from the WGA dataset were used for model validation. In this analysis, we used the Sequel II sequencing kit 2.0 on a PacBio Sequel II sequencer to obtain the WGA and M.SssI-treated DNA datasets for training and testing the HK model.

[0145] We used 2,000,000 CpG sites from M.SssI-treated DNA samples (fully methylated) and 2,000,000 CpG sites from WGA samples (fully unmethylated) to train the HK model. We used 2,000,000 CpG sites from M.SssI-treated DNA samples (fully methylated) and 2,000,000 CpG sites from WGA samples (fully unmethylated) to train the HK model EMA model. The optimal parameters of the HK or EMA model were obtained when the total prediction error between the output scores in the output layer calculated by the sigmoid function and the desired target output (binary value: 0 or 1) reached the minimum by iteratively adjusting the model parameters. In the deep learning algorithm, the total prediction error was measured by the sigmoid cross-entropy loss function. The model parameters learned from the training dataset were used to analyze the test dataset to output probability scores (called methylation scores), indicating the likelihood that the CpG site was methylated. The methylation scores output by the HK model were continuous probability scores from 0 to 1, rather than discrete binary values. For example, if the methylation score of a CpG site was 0.9, applying a methylation score threshold of 0.5 would result in the classification of this CpG site as methylated. Conversely, if the methylation score was 0.1, applying the same methylation score threshold of 0.5 would result in the classification of this site as unmethylated. Precision Comparison between C.HK Model and EMA Model

[0146] Figure 3A The receiver operating characteristic (ROC) curves comparing the HK and EMA models are shown. The x-axis shows specificity. The y-axis shows sensitivity. The dashed line shows the results of the HK model. The solid line shows the results of the EMA model. The area under the curve (AUC) of the EMA model was 0.98, which was improved compared to the HK model (AUC: 0.94) (P value < 0.0001, DeLong test).

[0147] Figure 3B A table showing the sensitivity of the HK and EMA models at a given specificity is shown. The first column lists different specificities. The second column lists the corresponding sensitivities of the HK model. The third column lists the corresponding sensitivities of the EMA model. At a specificity of 99%, the sensitivity given by the EMA model was 78%, which was better than the sensitivity of 58% of the HK model. At a specificity of 95%, the sensitivity given by the EMA model was 89%, which was better than the sensitivity of 76% of the HK model. At a specificity of 90%, the sensitivity given by the EMA model was 94%, which was better than the sensitivity of 85% of the HK model.

[0148] Figure 4ADisplays the receiver operating characteristic (ROC) curves for different models and datasets. The x-axis shows specificity. The y-axis shows sensitivity. The lowest curve shows HK Model 1 from a previously published study 1 HK Model 1 has an AUC of 0.91. The training dataset includes PCR-amplified DNA (i.e., unmethylated DNA; negative dataset) and the M.SssI-treated DNA set (i.e., methylated DNA; positive dataset), each involving 350,000 CpG sites.

[0149] The next lowest curve is labeled HK Model 2, which uses the same training dataset as a previously published study (HK Model 1) 1 (referred to as the PNAS dataset) 1 For the independent test dataset, the AUC is 0.97. Compared to HK Model 1, HK Model 2 is significantly improved.

[0150] The highest curve is HK Model 2 with a different training dataset. By preparing a new dataset according to the experimental protocol of Tse et al. (named New Dataset 01) 1 the training dataset size is increased to 13 million CpG sites. The performance of HK Model 2 is indeed further improved to an AUC of 0.99. If we define the cut-off value of the base modification score as 0.5, we can obtain 96% specificity and 95% sensitivity.

[0151] Figure 4B Is a graph showing precision and sub-read depth. The x-axis is the sub-read depth. The y-axis is the AUC. The bottom curve shows HK Model 1. The middle curve shows HK Model 2 using the PNAS dataset. The top curve shows HK Model 2 using New Dataset 01. A sub-read is defined as a read that starts at one adapter sequence and ends at another adapter sequence. In one embodiment, sub-reads that start or end in the middle of the inserted sequence can be used. The sub-read depth is defined as the number of sub-reads obtained from a strand of double-stranded DNA.

[0152] As Figure 4B shown, as the sub-read depth (x) increases, the AUC values of HK Model 2 and HK Model 1 gradually increase. For example, the sensitivity and specificity in HK Model 2 can reach 97% and 98% at a sub-read depth of >20x, while at a sub-read depth of 5 - 10x, the sensitivity and specificity are 87% and 89%. Notably, the AUC value of HK Model 2 trained with a large training dataset size (New Dataset 01) shows consistent improvement at different sub-read depths, indicating the robustness of HK Model 2. When comparing HK Model 2 (New Dataset 01) with HK Model 1 1 compared 1 1, the AUC of low read depths (1 - 5x) increased by 22%, and that of high read depths (>120x) increased by 6%. D. Single-stranded model

[0153] The data of HK model 2 shown above relate to the combined Watson and Crick data (i.e., double-stranded HK model 2). We explored the performance of HK model 2 in analyzing single-stranded data. Analyzing strand-specific methylation patterns would expand the applicability of the models proposed in this study. For example, the strand-specific HK model enables a detailed study of DNA hemimethylation reportedly occurring at CTCF (CCCTC-binding factor) / cohesin binding sites and playing a role in driving chromatin assembly. 13 .

[0154] Figure 5 is the ROC curve of HK model 2 using only single strands. The x-axis is specificity. The y-axis is sensitivity. Figure 5 shows that, compared with the double-stranded HK model 2, using the new dataset 01, the single-stranded HK model 2 can still achieve an AUC of 0.97 without significantly reducing the model performance. Figure 5 Also shown is the single-stranded HK model 2 for the new dataset 02, which has a higher AUC of 0.98. The new dataset 02 was prepared with a different protocol from that of the new dataset 01. The protocol is described below.

[0155] Figure 6 Shows the AUC of methylation analysis of CpG sites at positions relative to the most proximal end of the sequencing fragments for two datasets. The x-axis shows the relative distance (nt) to the end of the nearest DNA fragment. The y-axis shows the AUC. The upper figure shows the results from the new dataset 01 (by protocol A). The lower figure shows the results from the new dataset 02 (by protocol B).

[0156] As Figure 6 visible in the upper figure of, the performance of analyzing CpGs close to the 3'-end is generally worse than that of those close to the 5'-end. The closer the distance to the 3'-end, the lower the AUC value. We hypothesized that the reduced performance of the single-stranded model for those CpG sites close to the 3'-end might be due to unmethylated cytosines that were introduced into those M.SssI-treated DNA molecules with blunt ends during DNA end repair. To confirm this possibility, we modified protocol A to protocol B, where the end repair step was performed before the M.SssI treatment step to generate another new training dataset, i.e., the new dataset 02.

[0157] Using the new dataset 02 was able to distinguish methylated and unmethylated cytosines with an AUC of 0.98, confirming that the new protocol B is effective. More importantly, as Figure 6As can be seen, the AUC difference between the proximal regions of the 5' and 3' ends shown in the single-stranded HK model 2 disappears.

[0158] Figure 18 A scheme showing the improved performance of the single-stranded model for methylation sites near the 3' end is presented. Hollow lollipops represent unmethylated CpG sites. Solid lollipops represent methylated CpG sites. Scheme A for generating the new dataset 01 is shown on the left branch. Scheme B for generating the new dataset 02 is shown on the right branch.

[0159] Scheme A starts with M.SssI treatment, which methylates all CpG sites. Then the genomic DNA is sonicated. Then damage repair and end repair are performed. However, this damage repair and end repair result in the addition of unmethylated CpG sites to the fragments.

[0160] Scheme B starts with sonication. Then damage repair and end repair are performed. Again, unmethylated CpG sites are added to the fragments. Then M.SssI treatment is performed, which methylates the unmethylated CpG sites. As a result, all CpG sites are methylated. Scheme B can be used to generate a training set for the single-stranded model. E. Distinguishing different methylation types

[0161] The methods described herein can not only detect the presence of methylation, but also distinguish different types of methylation. This method can involve a single model for determining the type of methylation present, rather than using multiple models to test for the presence of each specific type of methylation. Detection of different types of methylation can use a single-stranded model or a double-stranded model. 1. Distinguishing 5mC and 5hmC using a test dataset

[0162] A training dataset including 5hmC modification is prepared. In contrast to preparing 5mC modification at CpG sites using a single methyltransferase (M.SssI), there is currently no such methyltransferase whose enzymatic reaction end product is 5hmC. TET (ten-eleven translocation) proteins can catalyze the stepwise oxidation of 5mC to produce a combination of 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine. The proportion of these oxidized cytosines in the DNA mixture after TET treatment varies according to the incubation time.

[0163] It has been reported that DNA treated with TET2 for about 5 minutes may result in a large percentage of 5mC and 5hmC in the reaction products, with relatively small contributions from 5fC and 5caC 14Therefore, we used TET to process DNA that had been methylated by M.SssI for 5 minutes previously, obtaining a training dataset (named the TET-5xC dataset) approximating a mixture of 5mC and 5hmC modifications.

[0164] Figure 8A Shows the composition of DNA after TET treatment. The amounts of 5hmC, 5fC, and 5acC vary with the incubation time.

[0165] Figure 8B Shows the preparation of a 5hmC detection dataset (named Lig-5hmCG) using ligation.

[0166] Figure 8C Shows the analytical workflow for 5mC and 5hmC detection. Based on the TET-5xC and WGA-uC datasets, we established a model (5xC detector) for determining 5xC and uC modifications. Based on M.SssI-mC and Lig-5hmCG, we established a model (5hmC detector) for further decomposing 5xC into 5mC and 5hmC modifications.

[0167] Figure 9A Shows the ROC curves of the test datasets for the 5xC and 5hmC detectors. Sensitivity is shown on the y-axis. Specificity is shown on the x-axis. For the 5xC and 5hmC detectors, we analyzed a total of 18,040,000 CpG sites and 325,851 CpG sites, obtaining AUC values of 0.99 and 0.97, respectively. Using a base modification score cutoff of 0.5, 93% specificity and 94% sensitivity were obtained for differentiating uC and 5mC, while 85% specificity and 94% sensitivity were obtained between 5hmC and 5mC.

[0168] Figure 9B Shows a box plot of the modification scores predicted by the 5hmC detector in the test dataset. The x-axis shows 5mC or 5hmC modification. The y-axis shows the predicted modification score. The modification score of 5hmC (median: 0.95; IQR: 0.92 - 0.96) is much higher than that of 5mC (median: 0.06; IQR: 0.05 - 0.17) (P value < 0.0001, Mann-Whitney U test).

[0169] Figure 9A and Figure 9B Demonstrates that the model can accurately distinguish between 5mC and 5hmC modifications. 2. Distinguishing 5mC and 5hmC in Biological Samples

[0170] We further used biological samples to demonstrate the effectiveness of the 5hmC detector. Buffy coat DNA samples were obtained from healthy individuals, and commercially available brain DNA samples were obtained from EpigenTek. We used bisulfite sequencing (BS-seq) and Tet-assisted bisulfite sequencing (TAB-seq) to deduce the 5xC (approximate total level of 5mC and 5hmC) and 5hmC levels in the buffy coat and brain samples, where the sequencing depth of the haploid genome was at least approximately 6-fold.

[0171] Figure 10A Shows the methylation levels measured in buffy coat and brain samples between different target genomic regions by different methods. The y-axis shows the methylation level. The x-axis shows the target genomic region. CGI is CpG island. LINE is long interspersed nuclear element. LTR is long terminal repeat. The upper panel shows the buffy coat. The lower panel shows the brain sample.

[0172] Figure 10A Shows that it was found that compared with buffy coat samples (range: 1.19% - 14.33%), 5hmC modifications deduced by HK model 2 were enriched in the brain among CpG islands (CGIs), enhancers, promoters, and repetitive regions (e.g., LINE, LTR, and satellite), with levels ranging from 2.23% to 27.47%. This 5hmC pattern was highly consistent with the data shown in the TAB-seq results [range of 5hmC levels: 4.78% - 27.64% (brain) versus 2.04% - 9.78% (buffy coat)]. It was found that in the buffy coat and brain, the total levels of cytosine modifications were generally consistent between the measurements of HK model 2 (represented by 5xC) and BS-seq.

[0173] Figure 10B Shows the methylation levels predicted by HK model 2 in human brain samples around the transcription start site (TSS) locus. The x-axis shows the distance (bp) relative to the TSS locus. The y-axis shows the methylation level inferred by HK model 2. The upper line is 5xC methylation. The middle line is 5mC methylation. The lower line is 5hmC methylation.

[0174] Notably, a sharp decline in the 5hmC and 5mC levels around the TSS was clearly seen in the results of the brain deduced by HK model 2, and the level of 5hmC was always lower than that of 5mC. This pattern was generally consistent with previous reports 7 。

[0175] Figure 10CShows the correlation of 5xC levels in brain samples measured by the HK model 2 and BS-seq. The x-axis shows the 5xC level (%) around the TSS site measured by BS-seq. The y-axis shows the 5xC level (%) around the TSS site inferred by the HK model 2.

[0176] Figure 10D Shows the correlation of 5hmC levels (%) in brain samples measured by the HK model 2 and TAB-seq. The x-axis shows the 5hmC level (%) around the TSS site measured by TAB-seq. The y-axis shows the 5hmC level (%) around the TSS site inferred by the HK model 2.

[0177] Importantly, the 5xC and 5hmC levels among positions near the TSS analyzed by the HK model 2 are linearly correlated with those measured by BS-seq (Pearson’s r: 0.99; P-value < 0.0001) and TAB-seq (Pearson’s r: 0.96; P-value < 0.0001). 5mC methylation can be determined by 5xC methylation that is not identified as 5hmC. These results demonstrate that the model can distinguish 5mC and 5hmC methylation. 3. Detection of 6mA and 4mC

[0178] Compared with 5mC, some DNA modifications show sharper kinetic changes at modified sites such as 6mA or 4mC (Flusberg B A et al., Nat methods 2010; Clark T A et al., Nucleic Acids Res 2012). Previous methods for detecting those DNA modifications such as 6mA involved comparing the IPD values at adenine (A) sites from native DNA sequencing data with control IPD values from unmethylated whole-genome amplified DNA or pre-computed in silico IPD models 15 . It has been reported that fixed cut-off values for the IPD ratio used for 6mA detection will introduce false-positive detections, especially from genomic regions with high sequencing depth 16 . These false-positive detections and general precision issues are addressed using the embodiments described herein. Adopting the HK model 2 framework, we developed a new adaptive method that includes generating a specific training data set and correcting the kinetic patterns within the measurement window to achieve high-performance detection of 6mA and 4mC.

[0179] Figure 11ASchematic diagram for the preparation of unmethylated and methylated adenine datasets (i.e., uA and 6mA datasets). We applied whole-genome amplification in the presence of 6mdATP, such that almost all adenine sites in the amplified DNA molecules would be 6mA (named WGA-6mA dataset). The corresponding negative dataset (named WGA-uA dataset) can be obtained from whole-genome amplification using unmodified dNTPs. Generate a similar training dataset for 4mC.

[0180] Figure 11B Shows the IPD distribution in the uA and 6mA datasets. The y-axis shows IPD. The x-axis shows the adenine methylation status. The IPD values at 6mA sites are significantly higher than those at uA sites (median: 0.90 vs. 0.22; P-value < 0.0001), indicating successful introduction of 6mA into the amplified DNA.

[0181] Figure 11C Shows the ROC curves for 6mA detection based on HK model 2 and based only on the IPD metric. The x-axis shows specificity. The y-axis shows sensitivity. The 6mA detector has an AUC of 0.99, which is superior to the analysis based on the IPD values at A sites (AUC: 0.94).

[0182] Figure 11D Shows the false positive rates for 6mA detection based on HK model 2 and based only on the IPD metric. The x-axis shows the classifier for 6mA. The y-axis shows the false positive rate. If the cut-off value for the 6mA modification fraction is set to 0.5, the sensitivity and specificity are 96% and 98% respectively. The corresponding false positive rate for HK model 2 is 1.7%, which is much lower than the method based only on the IPD metric (10.4%).

[0183] Figure 11E Shows the 6mA methylation levels determined by HK model 2 in non-GATC and GATC contexts in Dam-treated DNA samples. The x-axis shows the type of sites (non-GATC and GATC). The y-axis shows the predicted 6mA methylation levels. We further verified the performance of the 6mA detector using DNA molecules that had been treated with Escherichia coli DNA adenine methyltransferase (Dam), which is known to efficiently add a methyl group (i.e., 6mA) to adenine in the sequence context of 5’-GATC-3’. Figure 11EIt was shown that 95.2% of the GATC motifs were determined to have 6mA modification, while only 0.87% of the adenine sites had 6mA modification in the non-GATC context. The results further confirmed the effectiveness of 6mA determination using the 6mA detector. In some instances, based on the position where methylation is transferred, DNA methyltransferases can be classified into two categories: exocyclic amino methyltransferases and endocyclic methyltransferases. Exocyclic amino methyltransferases transfer a methyl group to the N4 position of cytosine (4mC) or the N6 position of adenine (6mA), such as Dam and CcrM. Endocyclic methyltransferases methylate cytosine at the C5 position (5mC), such as Dcm (Wion et al. Nat. Rev. Microbiol. 2006; 4: 183-192; Kumar et al. Nucleic Acids Res. 2018; 46: 3429-3445; Chen et al. Nat Commun. 2022; 13: 1248). These variants can be included in the analysis based on HK model 2.

[0184] Figure 12A The IPD distributions in the uC and 4mC datasets are shown. The y-axis shows the IPD. The x-axis shows the cytosine methylation status. The IPD values at the 4mC sites were significantly higher than those at the uC sites (median: 0.54 vs. 0.18; P-value < 0.001), indicating successful introduction of 4mC into the amplified DNA. The increase in IPD associated with 4mC seemed to be lower than that of 6mA (median: 0.54 vs. 0.90).

[0185] Figure 12B The ROC curves for 4mC detection based on HK model 2 and based on the IPD metric alone are shown. The x-axis shows the specificity. The y-axis shows the sensitivity. The 4mC detector had an AUC of 0.98, which was superior to the analysis based on the IPD values at the C sites (AUC: 0.92). The AUC for 6mA classification from uA was 0.94, while the AUC for 4mC classification from uC was 0.92, indicating that directly classifying 4mC using the IPD values would be more challenging.

[0186] To evaluate the performance of genome-wide 6mA detection in biological samples, we applied HK model 2 to analyze microbial DNA (with an average coverage of 220-fold). It is known that in Escherichia coli (E. coli) and Salmonella enterica but not in Bacillus subtilis (B. subtills), Enterococcus faecalis (E. faecalis), Listeria monocytogenes (L. mono), and Staphylococcus aureus (S. aureus), the sequence motif GATC is characterized by 6mA modification 17,18 The 6mA methylation levels at GATC among various microorganisms were analyzed by HK model 2.

[0187] Figure 13A Shows the 6mA methylation levels determined by the HK model 2. The x-axis shows the types of microbial DNA. The y-axis shows the 6mA methylation levels at the predicted GATC sites. The predicted median 6mA methylation levels associated with the GATC motif are 95% in both Escherichia coli and Salmonella enterica, while they are 2%, 1%, 2%, and 2% for Bacillus subtilis, Enterococcus faecalis, Listeria monocytogenes, and Staphylococcus aureus, respectively. The results match the expected results excellently.

[0188] Figure 13B Shows the de novo motif analysis associated with 6mA modification. The x-axis shows the relative distance to the observed 6mA sites. The y-axis shows bits (a measure of entropy). Interestingly, in addition to the well-known GATC motif, the respective 6mA-associated characteristic motifs for Bacillus subtilis, Escherichia coli, Enterococcus faecalis, Listeria monocytogenes, Staphylococcus aureus, and Salmonella enterica were determined to be ACA(N)8TG, AAGA(N)5CTC, CCAA(N)7TTG, GCA(N)7TGC, TA(N)6TA, CAGAG, respectively, which also correspond to previous studies 17,18 . These results further indicate that the 6mA detector can serve as a useful tool for 6mA analysis in real biological samples.

[0189] We further evaluated the performance in two signal pattern scenarios: (1) regions with few modifications, called the sparse signal pattern; (2) regions with many modifications, called the dense signal pattern. In one embodiment, a single measurement window (e.g., 21-nt window size) can contain multiple modified sites, potentially causing confounding effects from adjacent modifications and resulting in a model that does not generalize well.

[0190] Figure 27 Illustrates the sparse and dense signal patterns with 6mA modification. The sparse signal pattern can be defined as a region containing but not limited to no more than 2, 3, 4, or 5 modifications per 100 nucleotides. The dense signal pattern can be defined as a region containing but not limited to more than 5, 6, 7, or 8 modifications per 100 nucleotides. In some embodiments, the multi-base modifications can include different types of base modifications, such as 4mC, oxoG, 5mC, or any other modification described herein.

[0191] An example of a sparse signal pattern can be generated using the M.SssI methyltransferase that methylates cytosine only in the CpG context. The frequency of CpG in the human genome is approximately 1 in 100 nucleotides. An example of a dense signal pattern can be generated using whole genome amplification in which adenine is methylated. All adenine sites in the amplified DNA product will be methylated, and the frequency of adenine in the human genome is approximately 30 in 100 nucleotides.

[0192] To be able to detect single or multiple 6mA sites present in a measurement window using a common model, we developed a new normalization strategy to make the pattern of multiple 6mA signals comparable to that of a single 6mA signal.

[0193] Figure 15A The distribution of kinetic features in the measurement window between before and after signal normalization in different datasets is shown. Figure 1510 shows single 6mA data in a biological sample (sparse signal pattern, with one 6mA site in the measurement window). Figure 1520 shows uA in the training dataset (with no 6mA sites in the measurement window). Figure 1530 shows 6mA in the training (dense signal pattern, with multiple 6mA sites in the measurement window). Figure 1530 and Figure 1510 have significantly different signal patterns. Thus, when analyzing biological samples, training on the dense signal pattern may lead to misclassification. To solve this problem, we developed a method for correcting kinetic signals in the measurement window. For this purpose, we applied a denoising process 1540 (referred to as a denoiser), but used the median of the thymine signals.

[0194] Figure 15B The density distribution of kinetic features in different bases of template DNA based on the PacBio SequelII kit 2.0 is shown. Three randomly sampled data (replicates 1, 2, and 3) from a WGA - uA dataset were analyzed. We observed that the distribution of kinetic values was similar between unmodified adenine and thymine (e.g., peaks 1580 and 1590).

[0195] For the denoising process 1540, all kinetic values in the measurement window are divided by the median kinetic value of thymine. Thus, the normalized kinetic signal of uA exhibits a distribution with a mean of 1. Since adjacent 6mA sites may confound the target site analysis, the kinetic values regarding those "confounding sites" are set to 1, similar to the uA distribution to minimize the confounding effect during training.

[0196] Figure 1550 shows the denoised single 6mA biological data. Figure 1560 shows the denoised uA training data. Figure 1570 shows the denoised 6mA training data. Figure 1570 is similar to Figure 1550 after denoising.

[0197] Figure 11B - 11E The 6mA detector was built using the HK model 2 trained with the normalized data from the WGA-6mA and WGA-uA datasets. As a result, the 6mA detector can achieve an AUC of 0.99. Figure 12A and 12B The 4mC detector was built using the HK model 2 trained with the normalized data from the WGA-4mC and WGA-uC datasets. The 4mC detector outperforms the conventional assay (AUC: 0.98 vs. 0.92).

[0198] Figure 16 Shows the performance of the 6mA classifier using different types of 6mA signals. The columns show different types of classifiers, including those using IPD and HK model 2. For HK model 2, measurement size windows of 7nt and 21nt were used. For each measurement window size, models trained with and without a denoiser were used. The rows show the results for classifying different signals. For each pair of columns, the first column shows the sensitivity and the second column shows the specificity. Row 1610 shows the results from dense signals. Row 1630 shows the results from sparse signals prepared from DNA treated with Dam methyltransferase. The sparse signal pattern prepared by Dam treatment introduces 6mA into the GATC motif.

[0199] As shown in row 1610, without using a denoiser, when the measurement window size is enlarged, the sensitivity and specificity of the HK model 2 (6mA detector) in dense signal data increase. Without using a denoiser, compared with the IPD-based method (sensitivity: 91.8%; specificity: 89.88%), the HK model 2 results in higher sensitivity and specificity in the 7-nt and 21-nt window sizes, 95.7% and 98.8%, 97.45% and 99.4% respectively.

[0200] However, for the sparse signal pattern prepared by Dam treatment, the sensitivity decreases by 69.38% and 15.05% in the 7-nt and 21-nt window sizes respectively. These results indicate that reducing the window size can partially alleviate the overfitting problem.

[0201] After integrating the denoiser into the HK model 2, robust performance has been achieved in terms of sensitivity and specificity across all test datasets. For example, in line 1610, for the dense signal patterns obtained from whole-genome amplification, in the data with a 21-nt window with the denoiser, the sensitivity and specificity are 96.87% and 98.04%, respectively. In line 1630, for the sparse signal patterns obtained from Dam treatment, in the data with a 21-nt window with the denoiser, the sensitivity and specificity are 95.02% and 99.13%, respectively. These test results indicate that the denoiser developed in the present disclosure can be used for the detection of 6mA with high adaptability and stability. 4. Detection of Other Methylations

[0202] As the Lig-5hmCG throughput increases, by preparing various training datasets mediated by DNA ligation, a multi-class model for distinguishing uC, 5mC, and 5hmC can be directly trained ( Figure 8B ). Since it has been reported that in genomic DNA, the abundances of 5fC and 5caC are 10 - 10,000 times lower than that of 5hmC among various tissues and cells examined, the actual impact of residual 5fC and 5caC on the analysis of real biological samples may be relatively small. The consistent results observed between 5xC and 5hmC patterns in brain tissues measured by the HK model 2, BS-seq, and TAB-seq at least partially support the existence of a small number of methylation types other than 5mC and 5hmC in genomic DNA. F. Application of the Model

[0203] We have demonstrated that the HK model 2 exhibits better performance in determining various types of base modifications and has multiple functions.

[0204] Figure 24 Shows different sensitivities at a given specificity for different double-stranded models. For all datasets, the HK model 2 outperforms the HK model 1 in all characteristics.

[0205] Figure 18 Shows different sensitivities at a given specificity for different single-stranded models. The HK model 2 can be used to detect different methylation types.

[0206] We next set out to investigate what the potential implications are for clinical and biological applications. Choy et al. recently demonstrated that based on the HK model 1, analyzing the methylation patterns of cfDNA molecules in patients with hepatocellular carcinoma (HCC) can detect HCC 3。Choy et al. introduced the HCC methylation score, which is obtained by comparing the methylation pattern of each long cfDNA molecule derived from HK model 1 with the methylation patterns of reference tissues (e.g., HCC tumor tissues and normal tissues). 3 。Using HK model 2, we re-analyzed the dataset of Choy et al. that included cfDNA molecules with 1 - 6 CpG sites and calculated the HCC methylation score.

[0207] Figure 19A Shows the HCC methylation scores determined by HK model 2 in healthy individuals, HBV carriers, and HCC patients using sequenced DNA molecules with 1 - 6 CpG sites. The x-axis shows the type of individuals. The x-axis shows the HCC methylation scores determined by HK model 2. We observed that the HCC methylation scores in HCC patients (median: 0.764; IQR: 0.751 - 0.802) were significantly higher than those in non-HCC individuals (i.e., healthy individuals and HBV carriers) (median: 0.733; IQR: 0.729 - 0.745) (P-value = 0.0001; Mann-Whitney U test). Figure 19A Shows that HCC patients have statistically significantly different methylation scores from non-HCC individuals.

[0208] The HCC methylation score is described in US2023 / 0279498A1, filed on November 23, 2022, which is incorporated herein by reference in its entirety for all purposes. Briefly, to assess an individual's risk of developing HCC, we adjusted the origin tissue analysis by comparing the methylation pattern of plasma DNA molecules with the methylation signature of HCC tumor tissues. The tissue methylomics of HCC tumor tissues was obtained from previous studies.

[0209] First, a score S(HCC) reflecting the similarity between the DNA molecule and the HCC tumor tissue is calculated by the following formula: where P j is the methylation status of CpG site j in the plasma DNA molecule, r j,HCC is the methylation index of the corresponding CpG site in the reference methylomics of the tumor tissue, and n is the total number of CpG sites in the plasma DNA molecule.

[0210] Similarly, another score S(BC) is calculated in the following way to determine the similarity of the methylation pattern between the DNA molecule and the buffy coat:

[0211] Finally, both S(HCC) and S(BC) are integrated into a metric called the HCC methylation score by the following formula: where T is the total number of plasma DNA molecules analyzed in an individual. The higher the HCC methylation score, the more likely the test sample has HCC.

[0212] Figure 19B The ROC curves are shown for classifying individuals with or without HCC using the HCC methylation score based on molecules with 1 - 6 CpG sites or at least 7 CpG sites. The x - axis is specificity. The y - axis is sensitivity. Compared to the HK model 1 (AUC: 0.75), the HCC methylation score based on the HK model 2 results in a higher AUC (0.91) in differentiating individuals with and without HCC. If we use a dataset including cfDNA molecules with at least 7 CpG sites, the performance of HCC detection can be further improved to 0.97. These results indicate that we can not only reproduce the previous results using the HK model 2, but also achieve better performance.

[0213] Another possible application of our 6mA detector is to infer nucleosome positioning. 6mA modification can be differentially introduced into chromatin via DNA adenine methyltransferases (e.g., Hia5) depending on its accessibility status 12 . Analyze the SMRT - seq results of human cell nuclei (K562 cell line) treated with Hia5 using the 6mA detector based on the HK model 2 12 . Then we identified the nucleosome positioning in the genomic regions near CTCF - binding sites (i.e., CCCTC - binding factor), which are known to be flanked by well - organized nucleosome patterns 12 . We analyzed the 6mA signals within 1 kb upstream and downstream relative to the CTCF - binding sites.

[0214] Figure 19C The patterns of 6mA levels in genomic loci relative to CTCF - binding sites are shown. The x - axis shows the relative distance to the CTCF - binding site. The y - axis shows the predicted 6mA methylation level.

[0215] As Figure 19CAs can be seen, the 6mA levels in genomic loci relative to CTCF binding sites show a periodic signal with an interval of approximately 180 bp, similar to the nucleosome array. We hypothesized that the distance between two consecutive peaks of 6mA levels could facilitate the determination of nucleosome positioning, and the magnitude of 6mA levels could indicate the openness of chromatin states. For example, higher methylation levels could indicate higher openness or lower protein occupancy. The nucleosome position can be determined by measuring the distance between two consecutive peaks. This application of methylation levels and binding sites was discussed in Stergachis et al., Science, Vol. 368, No. 6498 (2020), which is incorporated herein by reference in its entirety for all purposes. G. Discussion

[0216] We have developed an enhanced deep learning framework, named HK Model 2, for analyzing multi-base modifications of DNA molecules sequenced by SMRT-seq. The sensitivities of HK Model 2 for 5mC, 5hmC, and 6mA detection can reach 98%, 90%, and 99% respectively, and the overall specificity can exceed 90%. This framework has been implemented using a hybrid structure of deep learning models including CNN and transformers. CNN can effectively capture local feature patterns in the measurement window through the convolution process, and transformers can learn global feature patterns through the "self-attention" mechanism. 20 In addition, preparing the deep learning model can include training dataset preparation and data processing of input features (e.g., signal normalization).

[0217] In this study, we have designed solutions for training dataset preparation and signal processing, depending on the type of target base modification. For example, for 5mC detection, the DNA end repair process is rearranged before M.SsssI treatment to reduce the contamination of unmodified cytosines present in the training dataset of methylated DNA. The design of this experimental protocol has improved model performance, typically for those CpG sites close to the 3' end of DNA fragments. For 5hmC detection, we have developed a new DNA ligation-based method to obtain a training dataset with high purity of 5hmC modification. In addition, we use a unique signal normalization to extend the capabilities of HK Model 2 to 6mA detection to minimize the potential confounding effects of adjacent 6mA sites. Therefore, HK Model 2 can have equally good performance in detecting 6mA modifications, regardless of whether there are single or multiple 6mA modifications in the measurement window.

[0218] In addition to the performance evaluation of HK Model 2, we also explored the potential impact on applications using general and improved model frameworks. For example, long cfDNA has more CpG sites and carries enriched tissue-specific molecular information. 2,3, but typically has a relatively low sub-read depth. Due to the enhanced accuracy of 5mC detection for sequencing molecules with low sub-read depth, the tissue origin analysis of the recently identified long cfDNA molecules using HK model 2 should be superior to that using HK model 1. In fact, using HK model 2, the performance of HCC detection has been greatly improved to an AUC of 0.97. For the detection of 6mA, we accurately determined the 6mA modification in a true biological sample (i.e., microbial DNA), which usually sparsely contains a single 6mA site in the measurement. On the other hand, we can identify the jagged ends of cfDNA molecules in the nucleus and the accessibility characteristics of native chromatin fibers for the artificially introduced 6mA modification that usually occurs among many nearby positions. These data further demonstrate the robustness of the structure of HK model 2 developed in this study.

[0219] In summary, HK model 2 is a general and improved method for detecting multi-base modifications using single-molecule real-time sequencing, enhancing the current efforts in developing non-invasive cancer detection and carefully studying chromatin structure. Exemplary method for H. model training

[0220] Figure 20 is a flowchart of an exemplary process 2000 for detecting nucleotide methylation in a nucleic acid molecule. Process 2000 can be used to train a model to detect methylation. In some embodiments, Figure 20 one or more process blocks of can be performed by any system described herein, including system 2700. The methylation can be any methylation described herein, including 5mC (5-methylcytosine) or 6mA (N6-methyladenine).

[0221] At block 2010, a first plurality of first data structures are received. Each first data structure in the first plurality of data structures can correspond to a respective nucleotide window sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules. Each of the first nucleic acid molecules can be sequenced by measuring pulses in a signal corresponding to the nucleotide. The methylation can have a known first state in the nucleotide at a target position in each window of each first nucleic acid molecule. The known first state can be whether methylation is present or absent. Each first data structure can include values of one or more signal characteristics at positions within the respective window. Figure 2A and Figure 2B shows an example of a first data structure. The plurality of known first states can include known 5mC, 5hmC, 6mA, 4mC, 5fC, 5caC, 1mA, 3mA, 7mA, 3mC, 2mG, 6mg, 7mG, 3mT, and 4mT states.

[0222] The plurality of first nucleic acid molecules can include single-stranded nucleic acid molecules. The plurality of first nucleic acid molecules can be obtained by methylating sites after repairing damaged nucleic acid molecules subjected to sonication or after end-repairing nucleic acid molecules subjected to sonication (e.g., as Figure 18 described).

[0223] The signal can be an optical signal (e.g., fluorescence, chemiluminescence, or photometric signal) or an electrical signal. The optical signal can be from single molecule real-time sequencing. The signal can be generated by a nucleotide or a label associated with a nucleotide.

[0224] The electrical signal can be from nanopore sequencing. The electrical signal can be current, voltage, resistance, inductance, capacitance, or impedance. The electrical signal is described in U.S. Patent Publication No. 2022 / 0328135A1, filed Apr. 12, 2022, the entire content of which is incorporated herein by reference for all purposes.

[0225] Each window of each first data structure can include 4 or more consecutive nucleotides, including 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90, 100, 200, 500, or any range between any two numbers or including any two numbers, or more consecutive nucleotides. Each window can have the same number of consecutive nucleotides. For example, each window can include 21 consecutive nucleotides upstream of the nucleotide at the target position and 21 consecutive nucleotides downstream of the nucleotide at the target position. Each window can have a different number of consecutive nucleotides upstream of the nucleotide at the target position than downstream of the nucleotide at the target position.

[0226] The target position can be the center of the corresponding window. For a window spanning an even number of nucleotides, the target position can be the position immediately upstream or immediately downstream of the window center. In some embodiments, the target position can be at any other position of the corresponding window, including the first position or the last position. For example, if the window spans n nucleotides of a strand, from the first position to the nth position (upstream or downstream), the target position can be at any position from the first position to the nth position.

[0227] The windows can be overlapping. Each window can include nucleotides on the first strand of the first nucleic acid molecule and nucleotides on the second strand of the first nucleic acid molecule. The first data structure can also include values for the strand characteristics of each nucleotide within the window. The strand characteristic can indicate the presence of the nucleotide or the first or second strand. The window can include nucleotides in the second strand that are not complementary to the nucleotides at the corresponding positions in the first strand. In some embodiments, all nucleotides on the second strand are complementary to the nucleotides on the first strand. In some embodiments, each window can include nucleotides only on one strand of the first nucleic acid molecule.

[0228] One or more signal characteristics can include sequence context. One or more signal characteristics can include the identity of the nucleotide (e.g., A, T, C, or G) of each nucleotide within each window. For each nucleotide within each window, one or more signal characteristics can also include: the position of the nucleotide within the sample nucleic acid molecule, the width of the pulse corresponding to the nucleotide, and / or the inter-pulse duration (IPD) representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide.

[0229] Each data structure among the plurality of first data structures can exclude first nucleic acid molecules with an IPD or width below a cut-off value. For example, only first nucleic acid molecules with an IPD value greater than the 10th percentile (or the 1st, 5th, 15th, 20th, 30th, 40th, 50th, 60th, 70th, 80th, 90th, or 95th percentile) can be used. The percentile can be based on data from all nucleic acid molecules in one or more reference samples. The cut-off value for the width can also correspond to a percentile.

[0230] The width of the pulse can be the pulse width at half of the pulse maximum. The inter-pulse duration can be the time between the maximum of the pulse associated with the nucleotide and the maximum of the pulse associated with an adjacent nucleotide. The adjacent nucleotides can be neighboring nucleotides. The characteristics can also include the pulse height corresponding to each nucleotide within the window. The characteristics can also include the value of the strand characteristic, which indicates whether the nucleotide is on the first or second strand of the first nucleic acid molecule.

[0231] The position can be the nucleotide distance relative to the target position. When the nucleotide is one nucleotide away from the target position in one direction, the position can be +1, and when the nucleotide is one nucleotide away from the target position in the opposite direction, the position can be -1.

[0232] The signal characteristics can include a vector that includes first fragment statistical values of electrical signal fragments corresponding to nucleotides. The characteristics can include first regional statistical values of the electrical signal in a region of the nucleic acid molecule that is equal to or greater than the window.

[0233] The first set of statistical values can represent the average value of the electrical signal segments corresponding to the nucleotides. In some embodiments, the first set of statistical values can represent the variation of the electrical signals (e.g., standard deviation) of the electrical signal segments corresponding to the nucleotides. In an embodiment, the first set of statistical values can represent the normalized value of the average value of the electrical signal segments corresponding to the nucleotides. Normalization can include rescaling such that the first set of statistical values is within a certain range (e.g., a range from 0 to 1). Normalization can include using the median, average, and / or deviation of part or all of the nucleotide chain.

[0234] The vector can include a second set of statistical values representing the variation of the electrical signal segments corresponding to the nucleotides. The vector can include a third set of statistical values representing the normalized value of the first set of statistical values.

[0235] The first regional statistical value can represent the average or median value of the electrical signals in the region. In an embodiment, the first regional statistical value can represent the median or average value of the absolute value of the variation of the electrical signals from the average or median value of the electrical signals in the region. The variation can be the standard deviation. In some embodiments, the first regional statistical value can be optional.

[0236] The input data structure can further include a second regional statistical value representing the median or average value of the absolute value of the variation of the electrical signals from the average or median value of the electrical signals in the region.

[0237] The first plurality of first data structures can include 5,000 - 10,000, 10,000 - 50,000, 50,000 - 100,000, 100,000 - 200,000, 200,000 - 500,000, 500,000 - 1,000,000, or 1,000,000 or more first data structures. The plurality of first nucleic acid molecules can include at least 1,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000 or more nucleic acid molecules. As another example, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 sequence reads can be generated.

[0238] As described in the present disclosure, the plurality of first nucleic acid molecules can include molecules having aptamers. A subset of the plurality of first nucleic acid molecules can include extended nucleic acid molecules. The extended nucleic acid molecules can include sample nucleic acid molecules and aptamers. The aptamers can have known sequences. The corresponding nucleotide window can include at least one nucleotide in the aptamer. Details of using aptamers are discussed elsewhere in the present disclosure.

[0239] The first data structure can include sparse signal patterns (e.g., fewer than 1 methylation for 25 nucleotides). The methylation can be 6mA, 4mC, or any methylation described herein. The window can include 21 or fewer consecutive nucleotides. Each first data structure can include values of one or more signal characteristics corresponding to methylated nucleotides that do not exceed one nucleotide in the corresponding window. Each first nucleic acid molecule can include no more than one methylated nucleotide in any window corresponding to any of the first data structures in the first plurality of first data structures. For example, the plurality of first nucleic acid molecules can include first nucleic acid molecules having ends repaired with methylated nucleotides (e.g., as described in FIG. 30). In some embodiments, the plurality of first nucleic acid molecules can include first nucleic acid molecules treated with DNA adenine methyltransferase (Dam). In some embodiments, the enzyme for treatment can include exocyclic amino methyltransferases and endocyclic methyltransferases. Exocyclic amino methyltransferases can include Dam and CcrM. Endocyclic methyltransferases can include Dcm.

[0240] In some embodiments, the values of one or more signal characteristics of a portion of the plurality of first data structures can include one or more values corresponding to one or more nucleotides, the value being determined using signal characteristics measured for nucleotides other than the one or more nucleotides. For example, the one or more nucleotides can be adenine, and the nucleotides other than the one or more nucleotides are thymine. Statistical measures of one nucleotide type (e.g., mean, median, or mode) can be used to normalize other nucleotide types (as Figure 15A described by the denoiser of). Reducing the number of methylated nucleotides in the window can involve using statistical measures of signal characteristics in place of measures of the measured methylated nucleotides.

[0241] In block 2020, a plurality of first training samples are stored. Each first training sample can include one of the first plurality of first data structures and a first label indicating the first state of the nucleotide at the target position. The storing can be on any computer-readable medium described herein.

[0242] In block 2030, a model is trained. The model can be trained through blocks 2040 - 2080.

[0243] In block 2040, the plurality of first data structures are filtered through a convolutional layer to obtain a convolutional matrix. Block 2040 can correspond to Figure 1 stage 108 in. The corresponding convolutional matrix can have a lower dimension than the corresponding first data structure.

[0244] The convolutional layer can be part of a Convolutional Neural Network (CNN). The CNN can include a set of convolutional filters configured to filter a first plurality of first data structures. The filters can be any of the filters described herein. The number of filters per layer can be 10 - 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 150, 150 - 200 or more. The kernel size of the filters can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15 - 20, 20 - 30, 30 - 40 or more. The CNN can include an input layer configured to receive the filtered first plurality of filter data structures. The CNN can also include a plurality of hidden layers including a plurality of nodes. The first layer of the plurality of hidden layers can be coupled to the input layer.

[0245] In some embodiments, the model can include a Recurrent Neural Network (RNN). The RNN can replace the CNN and the results from the RNN can be used instead of the convolutional matrix.

[0246] At block 2050, a transformer layer is applied to the convolutional matrix to obtain a transformer matrix. Block 2050 can correspond to Figure 1 phase 112 in, and the transformer layer can be any of the transformer layers described herein. Applying the transformer layer can include generating a plurality of attention scores that quantify the correlations between positions of the convolutional matrix.

[0247] In some embodiments, block 2040 is optional and the first plurality of data structures can be fed directly to the transformer layer.

[0248] The transformer layer can include a plurality of parameters for training. These parameters can include the number of multi - head self - attention, QKV biases (queries, keys, and values in the attention mechanism), and the ratio of the multi - layer perceptron.

[0249] Generating the plurality of attention scores can include using a plurality of multi - head self - attention. The number of multi - head self - attention can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40 or 50, or in any range between any two numbers and including any two numbers.

[0250] K, Q, and V can be vectors filled with weights. The input can be multiplied by a set of weights of K, Q, and V. The multiplication can be performed many times, such as but not limited to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, 300, 400 or 500 times, or in a range between any two numbers and including any two numbers.

[0251] The application transformer layer may include a normalized convolutional matrix. A softmax function may be used to determine the attention scores.

[0252] At block 2060, methylation probabilities are generated using the transformer matrix. Generating the methylation probabilities may include applying one or more neural network layers to the transformer matrix. Applying one or more neural network layers may include performing weight multiplication or bias addition.

[0253] At block 2070, the methylation probabilities are used to determine the output. The output may be whether methylation is present. In some embodiments, the output may be "0" or "1" or other binary classification based on a comparison of the probability with a cutoff value. For example, a methylation probability greater than the cutoff value may result in an output of "1".

[0254] At block 2080, based on the output of whether the model matches or does not match the corresponding label of the first label when the first plurality of first data structures are input into the model, the parameters of the model are optimized using a plurality of first training samples. The output of the model specifies whether the nucleotide at the target position in the corresponding window has methylation. The parameters of the model may include a plurality of attention scores.

[0255] As part of training a machine learning model, the parameters of the machine learning model (such as weights, thresholds, e.g., those usable in activation functions in neural networks, etc.) may be optimized based on training samples (training set) to provide optimized accuracy when classifying nucleotide methylation at target positions. Various forms of optimization may be performed, such as backpropagation, empirical risk minimization, and structural risk minimization. A set of validation samples (data structures and labels) may be used to validate the accuracy of the model. Cross-validation may be performed using different portions of the training set for training and validation. The model may include a plurality of sub-models, thereby providing an overall model. The sub-models may be weaker models, which once combined provide a more accurate final model.

[0256] Process 2000 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in combination with one or more other processes described elsewhere herein.

[0257] Although Figure 20 exemplary blocks of process 2000 are shown, in some embodiments, process 2000 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in Figure 20 . Additionally or alternatively, two or more blocks of process 2000 may be performed in parallel. I. Exemplary Methylation Detection Method

[0258] Figure 21 is a flowchart of an exemplary process 2100 for detecting nucleotide methylation in nucleic acid molecules. In some embodiments, Figure 21 one or more process blocks of can be performed by system 2700 or any system described herein. Methylation can be any methylation described herein, including 5mC (5-methylcytosine), 6mA (N6-methyladenine), 5hmC, 4mC, 5fC, 5caC, 1mA, 3mA, 7mA, 3mC, 2mG, 6mg, 7mG, 3mT, or 4mT.

[0259] In block 2110, data obtained by sequencing a sample nucleic acid molecule by measuring pulses in a signal corresponding to a nucleotide is received. The signal can be an optical signal or an electrical signal. The signal can be any signal described herein, including process 2000. Values of one or more signal characteristics are obtained from the data. Method 2100 can include sequencing the sample nucleic acid molecule by any sequencing technique described herein. The sample nucleic acid molecule can be single-stranded or double-stranded.

[0260] In some embodiments, the data can be obtained by sequencing an extended nucleic acid molecule. The extended nucleic acid molecule can include the sample nucleic acid molecule and an aptamer. The aptamer can have a known sequence and can be any aptamer described herein.

[0261] One or more signal characteristics can include the nucleotide identity of each nucleotide within each window. For each nucleotide within each window, one or more signal characteristics can also include: the position of the nucleotide within the sample nucleic acid molecule, the width of the pulse corresponding to the nucleotide, and / or the inter-pulse duration representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide. One or more signal characteristics can include the sequence context. The signal characteristics can be any signal characteristics described herein. In some embodiments, the nucleotide window can include at least one nucleotide of the aptamer. The signal characteristics can be the signal characteristics described with respect to process 2000 or any signal characteristics described herein.

[0262] In block 2120, an input data structure is created. The input data structure can include a window of the nucleotides sequenced in the sample nucleic acid molecule. For each nucleotide within the window, the input data structure can include one or more values of one or more signal characteristics. The window of the input data structure can have characteristics similar to the window of each first data structure in process 2000.

[0263] The nucleotides within the window may or may not align with the reference genome. The nucleotides within the window can be determined using circular consensus sequence (CCS) without aligning the sequenced nucleotides to the reference genome. The nucleotides in each window can be identified by CCS without aligning to the reference genome. In some embodiments, the window can be determined without CCS and without aligning the sequenced nucleotides to the reference genome.

[0264] The nucleotides within the window can be enriched or filtered. Enrichment can be performed by methods involving Cas9. The Cas9 method can include using a Cas9 complex to cleave a double-stranded DNA molecule to form a cleaved double-stranded DNA molecule and ligating a hairpin adaptor to the ends of the cleaved double-stranded DNA molecule. Filtering can be performed by selecting double-stranded DNA molecules having a size within a size range. The nucleotides can be from these double-stranded DNA molecules. Other methods that preserve the methylation state of the molecule (e.g., methyl-binding proteins) can be used.

[0265] At block 2130, an input data structure is input into the model. The model can be trained by process 2000 or any method described herein. The model can include the framework described in Section I.A with respect to Figure 1 the framework described. The model can include one or more transformer layers, including any transformer layer described herein. In some embodiments, the transformer layer can generate a plurality of attention scores that quantify the correlations between positions of data in the input data structure. In some embodiments, the transformer layer can generate a plurality of attention scores that quantify the correlations between data positions in a convolutional matrix by filtering the input data structure through a convolutional layer.

[0266] At block 2140, the model is used to determine whether methylation is present in the nucleotides at the target position within the window in the input data structure.

[0267] In some embodiments, process 2100 can further include determining that the methylation type is a first type among multiple types (e.g., as described with respect to Figure 8A , 8B and 8C). Determining whether methylation is present can include determining the presence of methylation and determining that the methylation is the first type. For example, each type among the multiple types can be one of 5mC, 5hmC, 6mA, or any methylation described herein. In some embodiments, the determination of the methylation type can occur simultaneously with the determination of whether methylation is present. For example, the training set for the methylation type may be sufficient for the desired accuracy. The training set can include 1 million - 5 million, 5 million - 10 million, 10 million to 15 million, 15 million to 20 million, or more than 20 million methylation sites of a specific type.

[0268] The input data structure can be one of a plurality of input data structures. Each input data structure can correspond to a respective nucleotide window sequenced in a respective sample nucleic acid molecule of a plurality of sample nucleic acid molecules. The plurality of sample nucleic acid molecules can be obtained from a biological sample of an object. The biological sample can be any biological sample described herein. The process 2100 can be repeated for each input data structure. The method can include creating a plurality of input data structures. The plurality of input data structures can be input into a model. The model can be used to determine whether methylation is present in the nucleotide at the target position in the respective window of each input data structure. The plurality of sample nucleic acid molecules can be single-stranded, double-stranded, or a combination thereof.

[0269] Each sample nucleic acid molecule of the plurality of sample nucleic acid molecules can have a size greater than a cutoff size. For example, the cutoff size can be 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 500 kb, or 1 Mb. Having a size cutoff can result in a higher sub-read depth, both of which can improve the accuracy of methylation detection. In some embodiments, the method can include fractionating DNA molecules by size prior to sequencing the DNA molecules.

[0270] The plurality of sample nucleic acid molecules can be aligned with a plurality of genomic regions. For each genomic region of the plurality of genomic regions, the plurality of sample nucleic acid molecules can be aligned with the genomic region. The number of sample nucleic acid molecules can be greater than a cutoff number. The cutoff number can be a sub-read depth cutoff value. The sub-read depth cutoff number can be 1x, 10x, 30x, 40x, 50x, 60x, 70x, 80x, 900x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, or 800x, or any range between and including any of these numbers. The sub-read depth cutoff number can be determined to improve or optimize the precision. The sub-read depth cutoff number can be related to the number of the plurality of genomic regions. For example, the higher the sub-read depth cutoff number, the lower the number of the plurality of genomic regions.

[0271] It can be determined that methylation is present at one or more nucleotides. The presence of methylation at one or more nucleotides can be used to determine the classification of a disorder. The classification of the disorder can include using the number of methylations. The number of methylations can be compared to a threshold. This comparison can be used to determine whether a locus or region is hypermethylated or hypomethylated. Optionally or additionally, the classification can include the location of one or more methylations. The location of one or more methylations can be determined by aligning sequence reads of a nucleic acid molecule to a reference genome. If methylation is indicated at certain positions known to be associated with a disorder, then the disorder can be determined. For example, the pattern of methylation sites can be compared to a reference pattern of the disorder, and the disorder can be determined based on this comparison. A match or substantial match (e.g., 80%, 90%, or 95% or more) to the reference pattern can indicate the disorder or a high likelihood of the disorder. The disorder can be cancer or any disorder described herein (e.g., pregnancy-related disorders, autoimmune diseases).

[0272] A statistically significant number of nucleic acid molecules can be analyzed in order to provide an accurate determination of a disorder, tissue origin, or clinically relevant DNA portion. In some embodiments, at least 1,000 nucleic acid molecules are analyzed. In other embodiments, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 nucleic acid molecules or more can be analyzed. As another example, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 sequence reads can be generated.

[0273] The method can include determining that the classification of the disorder is that the subject has the disorder. The classification can include the level of the disorder classified using the number of methylations and / or methylation sites.

[0274] The presence of a clinically relevant DNA portion, fetal methylation signature, maternal methylation signature, presence of an imprinted gene region, origin tissue (e.g., from a sample containing a mixture of different cell types), or location of a CTCF binding site can be determined using the presence of methylation at one or more nucleotides. Whether a locus or region is hypermethylated or hypomethylated can indicate the origin of the fragment. For example, certain genomic loci are known to be hypermethylated in cell-free fetal DNA compared to cell-free maternal DNA. Clinically relevant DNA portions include, but are not limited to, fetal DNA portions, tumor DNA portions (e.g., from a sample containing a mixture of tumor cells and non-tumor cells), and transplant DNA portions (e.g., from a sample containing a mixture of donor cells and recipient cells).

[0275] The method may further include treating a disorder. Treatment can be provided based on the determined disorder level, the measured methylation, and / or the source tissue (e.g., of tumor cells isolated from the circulation of a cancer patient). For example, the identified methylation can be targeted with a specific drug or chemotherapy. The source tissue can be used to guide surgery or any other form of treatment. And, the disorder level can be used to determine how invasive any type of treatment will be.

[0276] Embodiments can include treating a disorder in a patient after determining the disorder level in the patient. Treatment can include any suitable therapy, drug, chemotherapy, radiation, or surgery, including any treatment described in the references mentioned herein. Information regarding treatment in the references is incorporated herein by reference.

[0277] Process 2100 can include additional embodiments, such as any single embodiment or any combination and / or combination of the embodiments described below and / or one or more other processes described elsewhere herein.

[0278] Although Figure 21 exemplary boxes of process 2100 are shown, in some embodiments, process 2100 can include additional boxes, fewer boxes, different boxes, or boxes arranged differently than Figure 21 those shown. Additionally or alternatively, two or more boxes of process 2100 can be executed in parallel. II. Using Kinetic Signals from Adapter Sequences

[0279] Sequence data generated from PacBio SMRT-seq is typically trimmed to remove ligated adapter sequences prior to downstream analysis. The measurement windows described herein can have a minimum size and can be analyzed using a number of nucleotides upstream and downstream of the target nucleotide. Target nucleotide sites (e.g., CpG sites) located at or near the ends of DNA molecules may not have sufficient upstream or downstream nucleotides to construct such measurement windows for kinetic signal data for methylation analysis. Thus, regions containing CpG sites near the ends of DNA fragments are typically classified as undetectable regions. On average, 20.5% of cell-free DNA molecules contain at least one CpG that falls within an undetectable region (i.e., within 10 nt of the end). Thus, the use of measurement windows can result in data loss or reduced informativeness of methylation analysis.

[0280] The embodiments described herein can use kinetic signals (e.g., IPDs and PWs) derived from trimmed adapter sequences to construct complete measurement windows for target nucleotides (e.g., CpG sites) near the ends of fragments. Thus, CpG sites near the ends that would otherwise be unanalyzable can be analyzed.

[0281] Figure 22 Shows a plot of detectable CpG sites relative to the distance to the nearest end. The x-axis is the relative distance to the nearest DNA fragment end in base pairs. The y-axis is the percentage of detectable CpG sites. Figure 22 Shows a rapid decrease in the percentage of detectable CpG sites near the fragment end within a nucleotide distance of 11nt using HK model 1. The grey rectangle represents the non-detection region of HK model 1. To overcome this problem, HK model 2 utilizes kinetic signals retrieved from the sequencing adaptor to facilitate methylation analysis of CpGs near the fragment end. The percentage of detectable CpG sites in HK model 2 rebounds to almost 100%.

[0282] Figure 23 Shows workflow 2300 for analyzing site methylation using an adaptor sequence.

[0283] At box 2304, locate the interface between the inserted human DNA and the adaptor in the circularized template DNA. Known adaptor sequences can be identified using pairwise alignment.

[0284] At box 2308, retrieve data from adjacent ends of the adaptor and the inserted human DNA. The data can include kinetic signals (e.g., IPD, PW) and the identity of the nucleotides. These kinetic signals can be any of the kinetic features described herein, including Sections I.A.1 and II.A.

[0285] At box 2312, train a model for determining methylation of sites near the fragment end. The training model can include a training dataset having target nucleotides at or near the end (e.g., 10nt from the end). The training dataset can also include target nucleotides away from the end (e.g., 10nt away from the nearest end).

[0286] The trained model can then be used to analyze methylation of sites near the fragment end. In some embodiments, the trained model can be dedicated to nucleotides near or at the end of a nucleic acid molecule. The trained model can be used only after comparing the position of the target nucleotide to a threshold (e.g., 10nt from the end), and if the position is below the threshold, the trained model can be used.

[0287] Figure 24 Is a performance plot of the EMA model for determining the methylation status of CpG sites within 10nt of the 5’ end of a DNA fragment. The x-axis shows specificity. The y-axis shows sensitivity. The model using data from the adaptor sequence achieved an AUC of 0.97 for differentiating methylation and non-methylation of those CpG sites at a distance relative to the 5’ end of 10nt. A. Exemplary methods for model training

[0288] Figure 25 is a flowchart of an exemplary process 2500 for detecting nucleotide methylation in nucleic acid molecules. Process 2500 can be used to train a model using nucleic acid molecules with aptamers. In some embodiments, Figure 25 one or more process blocks of can be performed by system 2700 or any system described herein. The methylation can be any methylation described herein, including 5mC (5-methylcytosine) or 6mA (N6-methyladenine).

[0289] At block 2510, a first plurality of first data structures are received. Each first data structure in the first plurality of first data structures can correspond to a respective nucleotide window sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules. Each of the first nucleic acid molecules can be sequenced by measuring a pulse in a signal corresponding to the nucleotide. Each first nucleic acid molecule can include a training sample nucleic acid molecule and a first aptamer with a known sequence. In some embodiments, the known sequence can be the identity of a single nucleotide ligated to the end of the first nucleic acid molecule. The methylation can have a known first state in the nucleotide at a target position in a portion of each window of each first nucleic acid molecule corresponding to the training sample nucleic acid molecule. Each first data structure can include values of one or more signal characteristics. The plurality of known sequences of the first aptamers corresponding to the plurality of first nucleic acid molecules can be the same or different.

[0290] The signal can be an optical signal or an electrical signal. The optical signal can be from single molecule real-time sequencing. The electrical signal can be from nanopore sequencing. The signal can be any signal described herein, including process 2000.

[0291] One or more signal characteristics can include sequence context. One or more signal characteristics can include the nucleotide identity of each nucleotide within each window. For each nucleotide within each window, one or more signal characteristics can further include: the position of the nucleotide within the sample nucleic acid molecule, the width of the pulse corresponding to the nucleotide, and / or the interpulse duration representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide. The signal characteristics can be the signal characteristics described with respect to process 2000 or any signal characteristics described herein.

[0292] A subset of the windows can include at least 1, 2, 3, 4, 5, 6, 7, 8 or more nucleotides in the aptamer.

[0293] The nucleotides within each window can be determined using circular consensus sequences and without aligning the sequenced nucleotides to a reference genome.

[0294] Block 2510 can be performed similar to block 2010.

[0295] At block 2520, multiple first training samples are stored. Each first training sample can include one of a first plurality of first data structures and a first label indicating the first state of a nucleotide at a target position. Block 2520 can be performed similarly to block 2020.

[0296] At block 2530, the model is trained by using the multiple first training samples to optimize the parameters of the model based on the output of whether the model matches or does not match the corresponding label of the first label when the first plurality of first data structures are input into the model. The output of the model can specify whether the nucleotide at the target position in the corresponding window has methylation.

[0297] In some embodiments, the training sample nucleic acid molecule can have two adaptors. The known sequence is the first known sequence. Each first nucleic acid molecule can include a first adaptor at a first end. Each first nucleic acid molecule can include a second adaptor at a second end. The second adaptor can have a second known sequence. Each training sample nucleic acid molecule can have a first adaptor at one end and a second adaptor at the other end. A subset of the windows can include at least one nucleotide in the second adaptor.

[0298] In some embodiments, the training sample nucleic acid molecule can be limited to having a nucleotide at a target position within a certain distance from the nearest end of the first nucleic acid molecule. For example, the target position can be within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 10 - 15, or 15 - 20 nucleotides from the end. In some embodiments, the training sample nucleic acid molecule can include nucleotides that are not limited to certain positions (e.g., the molecule can have a target position greater than 10 nt from the end).

[0299] The model may include a Convolutional Neural Network (CNN). The CNN may include a set of convolutional filters configured to filter the first plurality of data structures and optionally the second plurality of data structures. The filters may be any of the filters described herein. The number of filters per layer may be 10 - 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 150, 150 - 200 or more. The kernel size of the filters may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15 - 20, 20 - 30, 30 - 40 or more. The CNN may include an input layer configured to receive the filtered first plurality of data structures and optionally the filtered second plurality of data structures. The CNN may also include a plurality of hidden layers including a plurality of nodes. The first layer of the plurality of hidden layers is coupled to the input layer. The CNN may also include an output layer coupled to the last layer of the plurality of hidden layers and configured to output an output data structure. The output data structure may include features.

[0300] In some embodiments, the model may include a Recurrent Neural Network (RNN). The RNN may replace the CNN. The model may include a supervised learning model. The supervised learning model may include different methods and algorithms, including analytical learning, artificial neural networks, backpropagation, boosting (meta - algorithm), Bayesian statistics, case - based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum description length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random fields, nearest neighbor algorithm, Probably Approximately Correct (PAC) learning, link wave descent rules, knowledge acquisition methodologies, symbolic machine learning algorithms, sub - symbolic machine learning algorithms, support vector machines, Minimal Complexity Machines (MCM), random forests, classifier ensembles, ordinal classification, data preprocessing, handling imbalanced datasets, statistical relational learning, or Proaftn - a multi - criterion classification algorithm. The model may be linear regression, logistic regression, deep recurrent neural networks (e.g., long short - term memory, LSTM), Bayesian classifiers, Hidden Markov Models (HMM), Linear Discriminant Analysis (LDA), k - means clustering, Density - Based Spatial Clustering of Applications with Noise (DBSCAN), random forest algorithms, support vector machines (SVM), or any of the models described herein.

[0301] As part of training a machine learning model, the parameters of the machine learning model (such as weights, thresholds, e.g., those that can be used for activation functions in neural networks, etc.) can be optimized based on training samples (training set) to provide optimized accuracy in classifying nucleotide methylation at a target location. Various forms of optimization can be performed, e.g., backpropagation, empirical risk minimization, and structural risk minimization. A set of validation samples (data structures and labels) can be used to validate the accuracy of the model. Cross-validation can be performed using different portions of the training set for training and validation. The model can include multiple sub-models, thereby providing an overall model. The sub-models can be weaker models that, once combined, provide a more accurate final model.

[0302] In some embodiments, the training can include one or more transformer layers and can be performed similar to blocks 2030 - 2080.

[0303] Process 2500 can include additional embodiments, such as any single embodiment described herein or any combination and / or incorporation of one or more other processes described elsewhere herein.

[0304] Although Figure 25 exemplary blocks of process 2500 are shown, in some embodiments, process 2500 can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown Figure 25 herein. Additionally or alternatively, two or more blocks of process 2500 can be performed in parallel. B. Exemplary Methods for Detecting Methylation

[0305] Figure 26 is a flowchart of an exemplary process 2600 for detecting nucleotide methylation in a nucleic acid molecule. In some embodiments, Figure 26 one or more process blocks of can be performed by system 2700 or any system described herein. The methylation can be any methylation described herein, including 5mC (5-methylcytosine) or 6mA (N6-methyladenine).

[0306] At block 2610, data obtained by sequencing an extended nucleic acid molecule by measuring pulses in a signal corresponding to a nucleotide is received. The signal can be an optical signal or an electrical signal. The extended nucleic acid molecule can include a sample nucleic acid molecule and an aptamer. The aptamer can have a known sequence. Values can be obtained from data of one or more signal characteristics. The signal characteristics can be any signal characteristics described herein, including those described with respect to process 2000 or block 2510. Block 2610 can be performed similar to block 2110.

[0307] In some embodiments, process 2600 may include attaching an aptamer to a sample nucleic acid molecule. In embodiments, the extended nucleic acid molecule may be sequenced using nanopore sequencing. In other embodiments, the extended nucleic acid molecule may be sequenced using single molecule real-time sequencing.

[0308] At block 2620, an input data structure is created. The input data structure may include a window of sequenced nucleotides in the extended nucleic acid molecule. The window may include at least one nucleotide in the aptamer. The window may include at least 1, 2, 3, 4, 5, 6, 7, 8 or more nucleotides in the aptamer. For each nucleotide within the window, the input data structure may include one or more values of one or more signal characteristics.

[0309] The nucleotides within the window may be determined using circular consensus sequence and without aligning the sequenced nucleotides to a reference genome. The window of the input data structure may have characteristics similar to the window of each first data structure in process 2000.

[0310] At block 2630, the input data structure is input into a model. The model may be trained by process 2500 or any method described herein. The known sequences may be the same as or different from the aptamer sequences in the training set. The model may include the framework described in Section I.A with respect to Figure 1 the framework described.

[0311] In some embodiments, the position of the target nucleotide may be determined and the distance from that position to the nearest end may be calculated. The distance may be compared to a threshold. If the distance is less than a certain threshold (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 10 - 15 or 15 - 20 nucleotides), the input data structure is input into the model. If the distance is greater than a certain threshold, the input data structure may be input into a second model that is not trained with a measurement window including nucleotides with aptamers.

[0312] At block 2640, it is determined using the model whether methylation is present in the nucleotide at the target position within the window in the input data structure. Block 2640 may be performed similar to block 2140.

[0313] In some embodiments, process 2600 may further include determining whether the methylation is more likely to be of a first type or a second type. For example, the first type may be one of 5mC, 5hmC, 6mA or any methylation described herein. The second type may be different from the first type. Process 2600 may not only determine the presence of methylation, but also the type of methylation present (e.g., as described with respect to Figure 8A , 8B and 8C).

[0314] As described with respect to process 2100, the input data structure can be one of a plurality of input data structures. The methylation assay can be used as described in process 2100.

[0315] In some embodiments, the extended nucleic acid molecule can include two aptamers. The aptamer is the first aptamer. The known sequence can be the first known sequence. The extended nucleic acid molecule can include the first aptamer at the first end. The extended nucleic acid molecule can include the second aptamer at the second end. The second aptamer can have a second known sequence. The window can include at least one nucleotide in the second aptamer.

[0316] Process 2600 can include additional embodiments, such as any single embodiment or any combination of embodiments as described herein and / or in combination with one or more other processes described elsewhere herein.

[0317] Although Figure 26 exemplary blocks of process 2600 are shown, in some embodiments, process 2600 can include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown Figure 26 herein. Additionally or alternatively, two or more blocks of process 2600 can be executed in parallel. III. Exemplary System

[0318] Figure 27 A measurement system 2700 according to an embodiment of the present disclosure is shown. The system shown includes a sample 2705, such as cell - free nucleic acid molecules (e.g., DNA and / or RNA) within an assay device 2710, where an assay 2708 can be performed on the sample 2705. For example, the sample 2705 can be contacted with the reagents of the assay 2708 to provide a signal (e.g., an intensity signal) of a physical characteristic 2715 (e.g., sequence information of the cell - free nucleic acid molecule). An example of an assay device can be a flow cell including assay probes and / or primers or a tube through which droplets move (where the droplets include the assay). Another example of an assay device is a sequencing device. The physical characteristic 2715 (e.g., fluorescence intensity, voltage, or current) from the sample is detected by a detector 2720. The detector 2720 can make measurements at intervals (e.g., periodic intervals) to obtain data points that make up a data signal. In one embodiment, an analog - to - digital converter converts the analog signal from the detector to digital form multiple times.

[0319] The measurement device 2710 and the detector 2720 can form a measurement system, for example, a sequencing system that performs sequencing according to the embodiments described herein. The data signal 2725 is sent from the detector 2720 to the logic system 2730. As an example, the data signal 2725 can be used to determine the sequence and / or location in a reference genome of a nucleic acid molecule (e.g., DNA and / or RNA). The data signal 2725 can include various simultaneous measurements, such as different colored fluorescent dyes of the sample 2705 or different electrical signals of different molecules, and thus the data signal 2725 can correspond to multiple signals. The data signal 2725 can be stored in the local memory 2735, the external memory 2740, or the storage device 2745. The measurement system can be composed of multiple measurement devices and detectors.

[0320] The logic system 2730 can be or can include a computer system, an ASIC, a microprocessor, a graphics processing unit (GPU), etc. It can also include a display (e.g., a monitor, an LED display, etc.) and a user input device (e.g., a mouse, a keyboard, a button, etc.) or be coupled thereto. The logic system 2730 and other components can be part of a stand-alone or network-connected computer system, or they can be directly attached to or incorporated into a device (e.g., a sequencing device) that includes the detector 2720 and / or the measurement device 2710. The logic system 2730 can also include software executed in the processor 2750. The logic system 2730 can include a computer-readable medium that stores instructions for controlling the measurement system 2700 to perform any of the methods described herein. For example, the logic system 2730 can provide commands to a system that includes the measurement device 2710 to perform sequencing or other physical operations. Such physical operations can be carried out in a specific order, for example, adding and removing reagents in a specific order. Such physical operations can be performed by a robotic system, for example, including a robotic arm, such as can be used to obtain a sample and perform measurements.

[0321] The measurement system 2700 can also include a treatment device 2760, which can provide treatment to an object. The treatment device 2760 can determine the treatment and / or be used to perform the treatment. Examples of such treatments can include surgery, radiotherapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplantation. The logic system 2730 can be connected to the treatment device 2760, for example, to provide the results of the methods described herein. The treatment device can receive inputs from other devices, such as imaging devices and user inputs (e.g., to control the treatment, such as control on a robotic system).

[0322] Any computer system mentioned herein can utilize any appropriate number of subsystems. Examples of such subsystems are in Figure 28is shown in computer system 10. In some embodiments, the computer system includes a single computer device, where the subsystem can be a component of the computer device. In other embodiments, the computer system can include multiple computer devices, each computer device being a subsystem with internal components. The computer system can include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0323] Figure 28 The subsystems shown therein are interconnected via system bus 75. Additional subsystems such as printer 74, keyboard 78, storage device 79, monitor 76 (e.g., display screen, such as an LED), etc. are shown coupled to display adapter 82. Peripheral devices and input / output (I / O) devices coupled to I / O controller 71 can be connected to the computer system in any number of ways known in the art, such as through input / output (I / O) ports 77 (e.g., USB, Lightning). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 1500 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of multiple instructions from system memory 72 or storage device 79 (e.g., fixed disk such as a hard disk drive or optical disk), as well as the exchange of information between subsystems. System memory 72 and / or storage device 79 can be embodied as a computer-readable medium. Another subsystem is data acquisition device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another component and can be output to the user.

[0324] The computer system can include multiple identical components or subsystems, for example, connected together via external interface 81, via an internal interface, or via a removable storage device that can be connected to and removed from one component to another. In some embodiments, the computer system, subsystem, or device can communicate via a network. In this case, one computer can be considered a client and another computer can be considered a server, where each computer can be part of the same computer system. The client and server can each include multiple systems, subsystems, or components.

[0325] Aspects of the embodiments may be implemented in hardware circuitry (e.g., an application specific integrated circuit or a field programmable gate array) and / or using computer software with a general purpose programmable processor, in a modular or integrated manner, in the form of control logic. As used herein, a processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processors on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will know and understand other ways and / or methods of implementing the embodiments of the present invention using hardware as well as combinations of hardware and software.

[0326] Any software component or functionality described in this application may be implemented as software code executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, using, for example, conventional or object-oriented techniques. The software code may be stored on a computer-readable medium as a series of instructions or commands for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard disk drive of an optical medium or such as a compact disc (CD) or a digital versatile disc (DVD) or a Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.

[0327] Such a program may also be encoded and transmitted using a carrier signal suitable for transmission via wired, optical, and / or wireless networks in accordance with various protocols, including the Internet. Thus, a computer-readable medium may be created using a data signal encoded with such a program. The computer-readable medium encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard disk drive, a CD, or an entire computer system), and may exist on or within different computer products in a system or network. The computer system may include a monitor, a printer, or other suitable display for providing any of the results mentioned herein to a user.

[0328] Any method described herein may be performed, in whole or in part, by a computer system including one or more processors configured to perform the steps. Accordingly, embodiments may be directed to a computer system configured to perform the steps of any method described herein, potentially having different components for performing the corresponding steps or groups of corresponding steps. Although presented in numbered steps, the steps of the methods herein may be performed simultaneously or at different times or in a different order logically possible. Additionally, portions of these steps may be used in conjunction with portions of other steps from other methods. Moreover, all or part of the steps may be optional. Additionally, any step of any method may be performed using modules, units, circuits, or other means of a system for performing those steps.

[0329] As will be apparent to those skilled in the art upon reading the present disclosure, each of the individual embodiments described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any one of several other embodiments without departing from the scope or spirit of the present disclosure.

[0330] The foregoing description of the exemplary embodiments of the present disclosure has been presented for purposes of illustration and description and has been set forth in order to provide a complete disclosure and description to one of ordinary skill in the art of how to make and use the embodiments of the present disclosure. It is not intended to be exhaustive or to limit the present disclosure to the precise forms described, nor are they intended to indicate that the experiments are all or the only experiments conducted. Although the present disclosure has been described in some detail for purposes of clarity of understanding by way of illustration and example, it will be readily apparent to one of ordinary skill in the art from the teachings of the present disclosure that certain changes and methylations can be made without departing from the spirit or scope of the appended claims.

[0331] Accordingly, the foregoing merely illustrates the principles of the invention. It is to be understood that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Moreover, all of the examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the disclosure and are not to be construed as being limited to these specifically recited examples and conditions. Additionally, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof are intended to cover both structural and functional equivalents thereof. Moreover, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed to perform the same function regardless of structure. Accordingly, the scope of the invention is not limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the invention are embodied by the appended claims.

[0332] Unless stated to the contrary, a statement of “a,” “an,” or “the” is intended to mean “one or more.” Unless stated to the contrary, the use of “or” is intended to mean “inclusive or” and not “exclusive or.” A reference to a “first” component does not necessarily require the provision of a second component. Additionally, unless expressly stated, a reference to a “first” or “second” component does not limit the referenced component to a particular position. The term “based on” is intended to mean “at least partially based on.”

[0333] Claims may be drafted to exclude any element that could be optional. Thus, this statement is intended to serve as a basis for the use of exclusive terms such as “solely,” “only,” etc. in connection with the recitation of claim elements or for the use of “negative” limitations.

[0334] Where a range of values is provided, it is to be understood that, unless the context clearly dictates otherwise, each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in the stated range and any other stated or intervening value in that stated range is encompassed within the embodiments of the present disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded from the range, and each range that includes either, neither, or both of the limits is also included in the present disclosure, subject to any specific excluded limit within the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the present disclosure.

[0335] All patents, patent applications, publications, and descriptions mentioned herein are hereby incorporated by reference in their entirety for all purposes as if each individual publication or patent were specifically and individually indicated to be incorporated by reference herein and are incorporated by reference to disclose and describe the methods and / or materials associated with the referenced publications. None are considered prior art. IV. References 1.Tse OYO,et al.Genome-wide detection of cytosine methylation bysingle molecule real-time sequencing.Proc Natl Acad Sci U S A118,(2021). 2. Yu SCY, et al. Single-molecule sequencing reveals a large population of long cell-free DNA molecules in maternal plasma. Proc Natl Acad Sci U SA 118, (2021). 3. Choy LYL, et al. Single-Molecule Sequencing Enables Long Cell-Free DNA Detection and Direct Methylation Analysis for Cancer Patients. Clin Chem 68, 1151 - 1163 (2022). 4. Yu SCY, et al. Comparison of Single Molecule, Real-Time Sequencing and Nanopore Sequencing for Analysis of the Size, End-Motif, and Tissue-of-Origin of Long Cell-Free DNA in Plasma. Clin Chem 69, 168 - 179 (2023). 5. Lau BT, et al. Single-molecule methylation profiles of cell-free DNA in cancer with nanopore sequencing. Genome Med 15, 33 (2023). 6. Pastor WA, et al. Genome-wide mapping of 5-hydroxymethylcytosine in embryonic stem cells. Nature 473, 394 - 397 (2011). 7. Wen L, et al. Whole-genome analysis of 5-hydroxymethylcytosine and 5-methylcytosine at base resolution in the human brain. Genome Biol 15, R49 (2014). 8. Cui XL, et al. A human tissue map of 5-hydroxymethylcytosines exhibits tissue specificity through gene and enhancer modulation. Nat Commun 11, 6161 (2020). 9. Li W, et al. 5-Hydroxymethylcytosine signatures in circulating cell-free DNA as diagnostic biomarkers for human cancers. Cell Res 27, 1243-1257 (2017). 10. Song CX, et al. 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages. Cell Res 27, 1231-1242 (2017). 11. Luo GZ, Blanco MA, Greer EL, He C, Shi Y. DNA N(6)-methyladenine: a new epigenetic mark in eukaryotes? Nat Rev Mol Cell Biol 16, 705-710 (2015). 12. Stergachis AB, Debo BM, Haugen E, Churchman LS, Stamatoyannopoulos JA. Single-molecule regulatory architectures captured by chromatin fiber sequencing. Science 368, 1449-1454 (2020). 13. Xu C, Corces VG. Nascent DNA methylome mapping reveals inheritance of hemimethylation at CTCF / cohesin sites. Science 359, 1166-1170 (2018). 14. Tamanaha E, Guan S, Marks K, Saleh L. Distributive Processing by the Iron(II) / alpha-Ketoglutarate-Dependent Catalytic Domains of the TET Enzymes Is Consistent with Epigenetic Roles for Oxidized 5-Methylcytosine Bases. J Am Chem Soc 138, 9345-9348 (2016). 15. Beaulaurier J, Schadt EE, Fang G. Deciphering bacterial epigenomes using modern sequencing technologies. Nat Rev Genet 20, 157-172 (2019). 16. Kong Y, et al. Critical assessment of DNA adenine methylation in eukaryotes using quantitative deconvolution. Science 375, 515-522 (2022). 17. McIntyre ABR, et al. Single-molecule sequencing detection of N6-methyladenine in microbial reference materials. Nat Commun 10, 579 (2019). 18. Yang S, Wang Y, Chen Y, Dai Q. MASQC: Next Generation Sequencing Assists Third Generation Sequencing for Quality Control in N6-Methyladenine DNA Identification. Front Genet 11, 269 (2020). 19. Jiang P, et al. Detection and characterization of jagged ends of double-stranded DNA in plasma. Genome Res 30, 1144-1153 (2020). 20. Vaswani A, et al. Attention is all you need. Advances in neural information processing systems 30, (2017). 21. Ito S, et al. Tet proteins can convert 5-methylcytosine to 5-formylcytosine and 5-carboxylcytosine. Science 333, 1300-1303 (2011).

Claims

1. A method for detecting methylation of nucleotides in a nucleic acid molecule, the method comprising: Receiving data obtained by sequencing a sample nucleic acid molecule by measuring pulses in a signal corresponding to a nucleotide of the sample nucleic acid molecule, and obtaining values of one or more signal characteristics from the data; Creating an input data structure that includes a window around a target position of a nucleotide sequenced in the sample nucleic acid molecule, wherein the input data structure includes, for each nucleotide within the window, one or more values of the one or more signal characteristics; Inputting the input data structure into a model, wherein the model is a machine learning model that is trained by: Receiving a first plurality of first data structures, each of the first plurality of first data structures corresponding to a respective window around a respective target position of a nucleotide sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring pulses in a signal corresponding to the nucleotide, wherein the methylation has a known first state in the nucleotide at the respective target position in each window of each first nucleic acid molecule, and each first data structure includes values of the same characteristics as the input data structure; Storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the nucleotide at the respective target position; Filtering the first plurality of first data structures through one or more convolutional layers to obtain a plurality of convolutional matrices; Applying a transformer layer to the plurality of convolutional matrices to obtain a transformer matrix, wherein applying the transformer layer to a convolutional matrix includes generating a plurality of attention scores that quantify the correlations between positions of the convolutional matrix; Using the transformer matrix to generate methylation probabilities at the respective target positions of the first plurality of first data structures; Determining an output using the methylation probabilities; and Using the plurality of first training samples to optimize parameters of the model based on an output of the model that matches or does not match the respective label of the first label when the first plurality of first data structures are input into the model, wherein the output of the model specifies whether the nucleotide at the respective target position in the respective window has the methylation, and wherein the parameters of the model include the plurality of attention scores; Using the model to determine whether the methylation exists in the nucleotide at the target position within the window in the input data structure.

2. The method according to claim 1, wherein the signal is an optical signal or an electrical signal.

3. The method according to claim 1, wherein the one or more signal characteristics include: For each nucleotide within the window: The identity of the nucleotide.

4. The method according to claim 3, wherein the one or more signal characteristics further include: For each nucleotide within the window: The position of the nucleotide within the sample nucleic acid molecule, The width of the pulse corresponding to the nucleotide, or The inter-pulse duration representing the time between a pulse corresponding to the nucleotide and a pulse corresponding to an adjacent nucleotide.

5. The method according to claim 1, wherein generating the plurality of attention scores includes using a plurality of multi-head self-attention.

6. The method according to claim 1, wherein generating the methylation probability includes applying one or more neural network layers to the transformer matrix.

7. The method according to claim 6, wherein applying the one or more neural network layers includes performing weight multiplication or bias addition.

8. The method according to claim 1, wherein the corresponding convolution result has a lower dimension than the corresponding first data structure.

9. The method according to claim 1, wherein the methylation is 5mC (5-methylcytosine).

10. The method according to claim 1, wherein the methylation is 6mA (N6-methyladenine).

11. The method according to claim 1, wherein the window contains 13 consecutive nucleotides.

12. The method according to claim 1, wherein the number of consecutive nucleotides upstream of the nucleotide at the target position in the window of the input data structure is different from the number of consecutive nucleotides downstream of the nucleotide at the target position.

13. The method according to claim 1, wherein the window of the input data structure includes 21 consecutive nucleotides upstream of the nucleotide at the target position and 21 consecutive nucleotides downstream of the nucleotide at the target position.

14. The method according to claim 1, wherein: the data is obtained by sequencing an extended nucleic acid molecule, the extended nucleic acid molecule contains the sample nucleic acid molecule and an adaptor, the adaptor has a known sequence, and the nucleotide window includes at least one nucleotide in the adaptor.

15. The method according to claim 1, wherein determining whether the methylation is present includes: determining the presence of the methylation; and determining that the methylation is a first type among multiple types.

16. The method according to claim 15, wherein each type among the multiple types is selected from: 5mC, 5hmC, and 6mA.

17. The method according to claim 1, wherein the sample nucleic acid molecule is single-stranded.

18. The method according to claim 1, wherein the plurality of first nucleic acid molecules includes single-stranded nucleic acid molecules.

19. A method for detecting methylation of nucleotides in a nucleic acid molecule, the method comprising: receiving a first plurality of first data structures, each first data structure of the first plurality of first data structures corresponding to a respective position window around a respective target position of a nucleotide sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring a pulse in a signal corresponding to the nucleotide, wherein the methylation has a known first state in the nucleotide at the respective target position in each window of each first nucleic acid molecule, and each first data structure includes values of one or more signal characteristics at positions within the respective window; Store a plurality of first training samples, each first training sample including one of a first plurality of first data structures and a first label indicating the first state of the nucleotide at the corresponding target position; and Train a model by: Filter the first plurality of first data structures through one or more convolutional layers to obtain a plurality of convolutional matrices, Apply a transformer layer to the plurality of convolutional matrices to obtain a transformer matrix, wherein applying the transformer layer to a convolutional matrix includes generating a plurality of attention scores quantifying the correlations between positions of the convolutional matrix, Using the transformer matrix, generate methylation probabilities at the corresponding target positions of the first plurality of first data structures, Determine an output using the methylation probabilities, and Using the plurality of first training samples, when the first plurality of first data structures are input into the model, optimize the parameters of the model based on the output of the model matching or not matching the corresponding labels of the first labels, wherein the output of the model specifies whether the nucleotide at the corresponding target position in the corresponding window has the methylation, and wherein the parameters of the model include the plurality of attention scores.

20. The method of claim 19, wherein the signal is an optical signal or an electrical signal.

21. The method of claim 19, wherein the one or more signal characteristics include: For each nucleotide within each window: The identity of the nucleotide.

22. The method of claim 21, wherein the one or more signal characteristics further include: For each nucleotide within each window: The position of the nucleotide within the first nucleic acid molecule, The width of the pulse corresponding to the nucleotide, or The inter-pulse duration representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide.

23. The method of claim 19, wherein generating the plurality of attention scores includes using a plurality of multi-head self-attention.

24. The method of claim 19, wherein generating the methylation probabilities includes applying one or more neural network layers to the transformer matrix.

25. The method of claim 24, wherein applying the one or more neural network layers includes performing weight multiplication or bias addition.

26. The method of claim 19, wherein the corresponding convolutional result has a lower dimension than the corresponding first data structure.

27. The method of claim 19, wherein the methylation is 5mC (5-methylcytosine).

28. The method of claim 19, wherein the methylation is 6mA (N6-methyladenine).

29. The method of claim 19, wherein the window includes 13 consecutive nucleotides.

30. The method of claim 19, wherein the number of consecutive nucleotides upstream of the nucleotide at the corresponding target position of each window of each first data structure of the first plurality of first data structures is different from the number of consecutive nucleotides downstream of the nucleotide at the corresponding target position.

31. The method according to claim 19, wherein each window of each first data structure of the first plurality of first data structures comprises 21 consecutive nucleotides upstream of the nucleotide at the corresponding target position and 21 consecutive nucleotides downstream of the nucleotide at the corresponding target position.

32. The method according to claim 19, wherein: a subset of the plurality of first nucleic acid molecules comprises extended nucleic acid molecules, the extended nucleic acid molecules comprise sample nucleic acid molecules and aptamers, the aptamers have known sequences, and the corresponding nucleotide window comprises at least one nucleotide in the aptamer.

33. The method according to claim 19, wherein the first plurality of first data structures comprises first data structures of nucleotides having a known first state, the known first state comprising known 5mC, 5hmC, and 6mA methylation.

34. The method according to claim 19, wherein the plurality of first nucleic acid molecules comprises single-stranded nucleic acid molecules.

35. The method according to claim 34, wherein the plurality of first nucleic acid molecules is obtained by methylating sites after repairing damage to sonicated nucleic acid molecules or after end-repairing the sonicated nucleic acid molecules.

36. The method according to claim 19, wherein each first data structure of the first plurality of first data structures comprises values of one or more signal characteristics corresponding to methylated nucleotides of no more than one nucleotide in the corresponding window.

37. The method according to claim 36, wherein each first nucleic acid molecule comprises no more than one methylated nucleotide in any window corresponding to any first data structure of the first plurality of first data structures.

38. The method according to claim 37, wherein the plurality of first nucleic acid molecules comprises first nucleic acid molecules treated with DNA adenine methyltransferase (Dam).

39. The method according to claim 36, further comprising: for at least a portion of the first plurality of first data structures: replacing values of one or more signal characteristics measured at one or more positions within the corresponding window with values of signal characteristics measured for different nucleotides.

40. The method according to claim 39, wherein: the one or more nucleotides are adenine, and the nucleotides other than the one or more nucleotides are thymine.

41. The method according to claim 36, wherein the methylation is 6mA or 4mC.

42. The method according to claim 36, wherein the window comprises 21 or fewer consecutive nucleotides.

43. A method for detecting methylation in a sample nucleic acid molecule, the method comprising: receiving data obtained by sequencing the extended nucleic acid molecule by measuring pulses in a signal corresponding to a nucleotide of the extended nucleic acid molecule, the extended nucleic acid molecule comprising the sample nucleic acid molecule and an aptamer having a known sequence, and obtaining values of one or more signal characteristics from the data; Create an input data structure that includes a window around a target position of a nucleotide sequenced in the extended nucleic acid molecule, where the window includes at least one nucleotide in the aptamer, and where the input data structure includes, for each nucleotide within the window, one or more values of the one or more signal characteristics; Input the input data structure into a model, where the model is trained by: Receiving a first plurality of first data structures, where each first data structure in the first plurality of first data structures corresponds to a respective window around a respective target position of a nucleotide sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, where each of the first nucleic acid molecules is sequenced by measuring a pulse in a signal corresponding to the nucleotide, where each first nucleic acid molecule includes a training sample nucleic acid molecule and a training first aptamer having the known sequence, and where the methylation has a known first state in the nucleotide at a respective target position in a portion of each window of each first nucleic acid molecule corresponding to the training sample nucleic acid molecule, and each first data structure includes values of the same characteristics as the input data structure; Storing a plurality of first training samples, each including one of the first plurality of first data structures and a first label indicating the first state of the nucleotide at the respective target position, and Using the plurality of first training samples, when the first plurality of first data structures are input into the model, optimizing the parameters of the model based on the output of the model matching or not matching the respective label of the first label, where the output of the model specifies whether the nucleotide at the respective target position in the respective window has the methylation; And Using the model, determining whether the methylation is present in the nucleotide at the target position within the window in the input data structure.

44. The method of claim 43, wherein the signal is an optical signal or an electrical signal.

45. The method of claim 43, wherein the one or more signal characteristics include: For each nucleotide within the window: The identity of the nucleotide.

46. The method of claim 45, wherein the one or more signal characteristics further include: For each nucleotide within the window: The position of the nucleotide relative to the target position within the window portion corresponding to the sample nucleic acid molecule, The width of the pulse corresponding to the nucleotide, or The inter-pulse duration representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide.

47. The method of claim 43, wherein: The aptamer is a first aptamer, The known sequence is a first known sequence, The extended nucleic acid molecule includes the first aptamer at a first end, The extended nucleic acid molecule includes a second aptamer at a second end, and The second aptamer has a second known sequence.

48. The method of claim 47, wherein the window includes at least one nucleotide in the second aptamer.

49. The method according to claim 43, wherein the window comprises at least three nucleotides in the aptamer.

50. The method according to claim 43, wherein the nucleotides within the window are determined using a circular consensus sequence, and the sequenced nucleotides are not aligned to a reference genome.

51. The method according to claim 43, wherein determining whether the methylation is present comprises: determining the presence of the methylation; and determining that the methylation is a first type among multiple types.

52. The method according to claim 51, wherein each type among the multiple types is selected from 5mC, 5hmC, and 6mA.

53. A method for detecting methylation of nucleotides in a nucleic acid molecule, the method comprising: receiving a first plurality of first data structures, each of the first plurality of first data structures corresponding to a respective window around a respective target position of a nucleotide sequenced in a respective nucleic acid molecule of a plurality of first nucleic acid molecules, wherein each of the first nucleic acid molecules is sequenced by measuring a pulse in a signal corresponding to the nucleotide, wherein each first nucleic acid molecule comprises a training sample nucleic acid molecule and a first aptamer having a known sequence, wherein the methylation has a known first state in the nucleotide at a respective target position in a portion of each window of each first nucleic acid molecule corresponding to the training sample nucleic acid molecule, and each first data structure comprises values of one or more signal characteristics; storing a plurality of first training samples, each comprising one of the first plurality of first data structures and a first label indicating the first state of the nucleotide at the respective target position, and training a model by optimizing parameters of the model using the plurality of first training samples, when the first plurality of first data structures are input into the model, based on an output of the model matching or not matching a respective label of the first label, wherein the output of the model specifies whether the nucleotide at the respective target position in the respective window has the methylation.

54. The method according to claim 53, wherein the signal is an optical signal or an electrical signal.

55. The method according to claim 53, wherein the one or more signal characteristics comprise: for each nucleotide within each window: the identity of the nucleotide.

56. The method according to claim 55, wherein the one or more signal characteristics further comprise: for each nucleotide within the window: the position of the nucleotide relative to the target position within a portion of the window corresponding to the training sample nucleic acid molecule, the width of the pulse corresponding to the nucleotide, or the inter-pulse duration representing the time between the pulse corresponding to the nucleotide and the pulse corresponding to an adjacent nucleotide.

57. The method according to claim 53, wherein: the known sequence is a first known sequence, each first nucleic acid molecule comprises the first aptamer at a first end, each first nucleic acid molecule comprises a second aptamer at a second end, and the second aptamer has a second known sequence.

58. The method according to claim 57, wherein the subset of windows comprises at least one nucleotide in the second aptamer.

59. The method according to claim 53, wherein the subset of windows comprises at least three nucleotides in the first aptamer.

60. The method according to claim 53, wherein nucleotides within each window are determined using a circular consensus sequence without aligning the sequenced nucleotides to a reference genome.

61. A method for detecting methylation of nucleotides in a nucleic acid molecule, the method comprising: receiving data obtained by sequencing a sample nucleic acid molecule by measuring pulses in a signal corresponding to a nucleotide of the sample nucleic acid molecule; obtaining values of one or more signal characteristics from the data; creating an input data structure comprising windows around a target position of a nucleotide sequenced in the sample nucleic acid molecule, wherein the input data structure comprises, for each nucleotide within the window, one or more values of the one or more signal characteristics; inputting the input data structure into a model, wherein the model comprises a transformer layer that generates a plurality of attention scores that quantify the correlation between positions of data in the input data structure or quantify the correlation between positions of data in a convolution result from filtering the input data structure by a convolutional layer; and using the model to determine whether the methylation is present in the nucleotide at the target position within the window in the input data structure.

62. The method according to claim 61, wherein the model is trained by the method according to any one of claims 19-33.

63. The method according to any one of claims 1, 19, and 61, wherein the input data structure comprises a signal input matrix, and wherein different positions within the window correspond to different columns in the signal input matrix.

64. The method according to claim 63, wherein the plurality of attention scores form an attention matrix, and wherein generating the attention scores comprises: generating a query matrix Q using the signal input matrix and a query weight matrix, wherein the query matrix Q has the same number of columns as the signal input matrix; generating a key matrix K using the signal input matrix and a key weight matrix, wherein the key matrix K has the same number of columns as the signal input matrix; and applying each column of the query matrix Q to each column of the key matrix K to obtain the attention matrix.

65. The method according to any one of claims 63 and 64, wherein using the model comprises: generating a value matrix V using the signal input matrix and a value weight matrix, wherein the value matrix V has the same number of columns as the signal input matrix; and generating self-attention scores forming a self-attention matrix by applying the attention matrix to the value matrix V.

66. The method according to any one of claims 63-65, wherein using the model comprises: Combine the positional index matrix and the signal input matrix to obtain a transformer index matrix for generating the query matrix Q and the key matrix K.

67. The method according to any one of claims 1, 19, and 61, wherein using the model comprises: Filter the input data structure through one or more convolutional matrices to obtain one or more convolutional matrices; And Combine the one or more convolutional matrices to obtain a convolutional signal matrix, wherein different positions within the window correspond to different columns in the convolutional signal matrix.

68. The method according to claim 67, wherein the filtering uses a filter convolution kernel having a filter convolution kernel size smaller than the window size.

69. The method according to claim 67, wherein the plurality of attention scores form an attention matrix, and wherein generating the attention scores comprises: Generate a query matrix Q using the convolutional signal matrix and a query weight matrix, wherein the query matrix Q has the same number of columns as the convolutional signal matrix; Generate a key matrix K using the convolutional signal matrix and a key weight matrix, wherein the key matrix K has the same number of columns as the convolutional signal matrix; And Apply each column of the query matrix Q to each column of the key matrix K to obtain the attention matrix.

70. The method according to any one of claims 67 and 69, wherein using the model comprises: Generate a value matrix V using the convolutional signal matrix and a value weight matrix, wherein the value matrix V has the same number of columns as the convolutional signal matrix; And Generate self-attention scores forming a self-attention matrix by applying the attention matrix to the value matrix V.

71. The method according to any one of claims 67-70, wherein using the model comprises: Combine the positional index matrix and the convolutional signal matrix to obtain a transformer index matrix for generating the query matrix Q and the key matrix K.

72. The method according to any one of claims 63-71, wherein using the model comprises: Generate multiple sets of attention scores, each set of attention scores forming a corresponding attention matrix of a plurality of attention matrices; And Use the plurality of attention matrices to determine whether methylation is present in the nucleotide at the target position.

73. The method according to claim 72, wherein using the model comprises: Combine the plurality of attention matrices into a single output layer for determining whether methylation is present in the nucleotide at the target position.

74. The method according to any one of claims 63-73, wherein using the model comprises: Apply one or more fully connected layers to the output of the transformer layer.

75. The method according to claim 61, wherein determining whether methylation is present comprises: Determine the presence of the methylation; And Determine that the methylation is the first type among multiple types.

76. The method according to claim 75, wherein each type among the multiple types is selected from 5mC, 5hmC, and 6mA.

77. A method for detecting methylation in a sample nucleic acid molecule, the method comprising: Receiving data obtained by sequencing an extended nucleic acid molecule by measuring pulses in a signal corresponding to nucleotides of the extended nucleic acid molecule, the extended nucleic acid molecule including the sample nucleic acid molecule and an adaptor having a known sequence; Obtaining values of one or more signal characteristics from the data; Creating an input data structure including a window around a target position of a nucleotide sequenced in the extended nucleic acid molecule, wherein the window includes at least one nucleotide in the adaptor, and wherein the input data structure includes, for each nucleotide within the window, one or more values of the one or more signal characteristics; Inputting the input data structure into a model; And Using the model to determine whether the methylation is present in the nucleotide at the target position within the window in the input data structure.

78. The method according to claim 77, wherein determining whether the methylation is present comprises: Determining that the methylation is present; And Determining that the methylation is a first type among multiple types.

79. The method according to claim 78, wherein each type among the multiple types is selected from 5mC, 5hmC, and 6mA.

80. The method according to claim 77, wherein the model is trained by the method according to any one of claims 53 - 62.

81. A computer product comprising a non - transitory computer - readable medium storing multiple instructions, the multiple instructions when executed controlling a computer system to perform the method according to any one of the foregoing claims.

82. A system comprising: The computer product according to claim 81; And One or more processors for executing the instructions stored on the computer - readable medium.

83. A system comprising means for performing any of the above - described methods.

84. A system comprising one or more processors configured to perform any of the above - described methods.

85. A system comprising modules that respectively perform the steps of any one of the above - described methods.

Citation Information

Patent Citations

  • Determination of base modifications of nucleic acids

    US20210047679A1

  • Base modification analysis using electrical signals

    US20220328135A1

  • Molecular analyses using long cell-free DNA molecules for disease classification

    US20230279498A1