Methods, systems, and kits for sequencing nucleic acids with deamination and mutagenesis

By introducing mutations using methylation-specific deamination agents and sequencing based on mutation patterns, the method addresses the challenge of sequencing difficult nucleic acid sequences and measuring methylation levels, enhancing sequence assembly and haplotype identification.

WO2026072259A1PCT designated stage Publication Date: 2026-04-02ILLUMINA INC
View PDF 15 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing technologies struggle to accurately determine the sequence of difficult-to-distinguish sequences, such as repetitive regions or haplotypes, and to measure methylation levels in nucleic acid samples.

Method used

Introduce mutations into nucleic acids using methylation-specific deamination agents to convert methylated or unmethylated cytosines to thymine, followed by sequencing and assembling sequence reads based on patterns of mutations and deaminations to determine the nucleic acid sequence, and optionally phase sequence reads to identify haplotype-specific methylation levels.

Benefits of technology

Enhances the accuracy of nucleic acid sequencing by improving sequence assembly and enabling the identification of sequence variants and haplotype-specific methylation levels, particularly in regions with high similarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044591_02042026_PF_FP_ABST
    Figure US2025044591_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods, kits, and systems for determining a sequence of a nucleic acid. In some embodiments, the methods include contacting nucleic acids with a methylation-specific deamination agent, to form deaminated nucleic acids; introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids; amplifying and fragmenting the mutated and deaminated nucleic acids; sequencing the mutated and deaminated nucleic acids to obtain first sequence reads; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence. Further disclosed herein are methods of phasing sequence reads and methods of determining a haplotype-specific methylation level.
Need to check novelty before this filing date? Find Prior Art

Description

ILLINC 843WO / IP-2808-PCT PATENTMETHODS, SYSTEMS, AND KITS FOR SEQUENCING NUCLEIC ACIDS WITH DEAMINATION AND MUTAGENESISCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Prov. App. No. 63 / 701410 filed September 30, 2024 entitled “METHODS, SYSTEMS, AND KITS FOR SEQUENCING NUCLEIC ACIDS WITH DEAMINATION AND MUTAGENESIS” which is incorporated by reference herein in its entirety.FIELD

[0002] The present disclosure relates to the field of nucleic acid sequencing. More particularly, the present disclosure relates to the field of Sequencing Aided by Mutagenesis (SAM).BACKGROUND

[0003] Sequencing Aided by Mutagenesis (SAM) is a useful technique for determining the sequence of a nucleic acid, where mutations are intentionally introduced into sample nucleic acids, which are then sequenced, and the pattern of mutations is used to aid sequence assembly. SAM is particularly useful in the context of determining the sequence of nucleic acid samples which include sequences which are difficult to distinguish from one another, such as repetitive regions or haplotypes.

[0004] Methods for specifically converting methylated or unmethylated cytosines to thymine may be used to measure methylation in a nucleic acid sample.SUMMARY

[0005] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, andparticularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0006] Disclosed herein are methods for determining a sequence of a nucleic acid. In some embodiments, the method includes contacting nucleic acids with a methylationspecific deamination agent, to form deaminated nucleic acids; introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids; amplifying and fragmenting the mutated and deaminated nucleic acids; sequencing the mutated and deaminated nucleic acids to obtain first sequence reads; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

[0007] In some embodiments, the methylation-specific deamination agent comprises a deamination agent specific for 5-methyl cytosine (5mC) and 5- hydroxymethylcytosine (5hmC), or an unmethylated cytosine (uC) -specific deamination agent.

[0008] In some embodiments, assembling the first sequence reads comprises comparing the first sequence reads with a reference sequence or with second sequence reads to identify likely mutated bases or deaminated bases. In some embodiments, the second sequence reads comprise sequence reads from deaminated unmutated nucleic acids, or from unmutated nucleic acids that are not deaminated.

[0009] In some embodiments, assembling the first sequence reads comprises: grouping first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids, to generate groups of first sequence reads; and assembling each group of first sequence reads. In some embodiments, the method further comprises aligning assembled first sequence reads to a reference sequence or to second sequence reads, wherein the second sequence reads comprise sequence reads from deaminated unmutated nucleic acids, or from unmutated nucleic acids that are not deaminated.

[0010] In some embodiments, the method comprises assigning first sequence reads, groups of first sequence reads, or assembled first sequence reads to a strand based on nucleotide composition. In some embodiments, the method comprises assigning assembled first sequencereads to a strand based on cytosine (C) to thymine (T) and guanine (G) to adenine (A) imbalance arising from deamination of 5mC and 5hmC to T or deamination of uC to T.

[0011] In some embodiments, the method further comprises replacing a likely mutated base with a base from a reference sequence, a base from deaminated unmutated sequence reads, or from unmutated nucleic acids that are not deaminated.

[0012] In some embodiments, the method further comprises identifying a sequence variant in the nucleic acid sequence. In some embodiments, the sequence variant is identified as being associated with a specific strand of a double-stranded nucleic acid.

[0013] In some embodiments, the method further comprises phasing the first sequence reads. In some embodiments, the method comprises phasing the first sequence reads using 5-base phasing with 5mC and 5hmC, or uC, as a fifth base.

[0014] Disclosed herein are methods of determining a haplotype-specific methylation level. In some embodiments, the method includes phasing first sequence reads according to the methods described herein; and determining a haplotype-specific methylation level.

[0015] In some embodiments, the deamination agent specific for 5mC and 5hmC comprises a deaminase specific for 5mC and 5hmC. In some embodiments, the deaminase specific for 5mC and 5hmC comprises an engineered APOBEC enzyme or a SEM-seq deaminase. In some embodiments, the deamination agent specific for 5mC and 5hmC comprises TAPS reagents. In some embodiments, the uC-specific deamination agent comprises sodium bisulfite, or a uC-specific deaminase. In some embodiments, the uC-specific deamination agent comprises an oxidation enhancer and an APOBEC enzyme that is specific for non-oxidized cytosines.

[0016] In some embodiments, the method further comprises denaturing double stranded nucleic acids using heat, a denaturing agent, or a combination thereof.

[0017] In some embodiments, the method further comprises fragmenting nucleic acids and adding an adapter sequence to an end of the nucleic acids prior to contacting nucleic acids with a methylation-specific deamination agent. In some embodiments, fragmenting nucleic acids and adding an adapter sequence to an end of the nucleic acids comprises tagmenting nucleic acids using a transposome complex.

[0018] In some embodiments, introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids comprises amplifying the deaminated nucleic acids with a polymerase, dNTPs, and a nucleotide analog.

[0019] Disclosed herein are kits for determining a sequence of a nucleic acid. In some embodiments, the kit includes a methylation-specific deamination agent, and one or more mutagenesis agents.

[0020] Further disclosed herein are computer-implemented methods for determining a sequence of a nucleic acid. In some embodiments, the method includes receiving first sequence reads from mutated and deaminated nucleic acids; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

[0021] Further disclosed herein are systems for determining a sequence of a nucleic acid. In some embodiments, the system includes one or more processors having instructions that when executed perform a method comprising: receiving first sequence reads from mutated and deaminated nucleic acids; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

[0022] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive first sequence reads from mutated and deaminated nucleic acids; and assemble the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the followingdescriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.

[0024] FIG. 1 is a flow diagram that schematically illustrates an exemplary method for determining a sequence of a nucleic acid.

[0025] FIG. 2A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.

[0026] FIG. 2B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 2A.

[0027] FIG. 3 schematically illustrates an exemplary workflow.

[0028] FIG. 4 schematically illustrates two options for analysis of assembled first sequence reads.

[0029] FIG. 5 is a graph illustrating strand inference accuracy with respect to read length.

[0030] FIG. 6 is a graph illustrating median phase block size with respect to minimum usable methylation imbalance, for various sequence read fragment sizes.DETAILED DESCRIPTION

[0031] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.

[0032] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporatedherein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0033] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.

[0034] Disclosed herein are methods and kits for determining a sequence of a nucleic acid. In some embodiments, the methods include contacting nucleic acids with a methylation-specific deamination agent, to form deaminated nucleic acids. In some embodiments, the methylation-specific deamination agent can specifically convert methylated cytosine bases into thymine bases. In other embodiments, the methylation-specific deamination agent specifically converts unmethylated cytosine bases to thymine bases.

[0035] Next, mutations are introduced into the deaminated nucleic acids to generate mutated and deaminated nucleic acids. A variety of techniques may be used for random mutagenesis, for example, amplification with a nucleotide analog which can be incorporated into the deaminated nucleic acids to provide transition mutations. In some embodiments, the mutations are random transition mutations. A sequencing library can be prepared from the mutated and deaminated nucleic acids, for example by amplifying the mutated and deaminated nucleic acids to create multiple copies, and fragmenting the mutated and deaminated nucleic acids into smaller nucleic acid molecules that are more amenable to some sequencing technologies. Next, the mutated and deaminated nucleic acids can be sequenced, such as in a sequencing by synthesis (SBS) system, to obtain first sequence reads.

[0036] The first sequence reads can be assembled into a final sequence based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence. The sequence assembly step can reconstruct the longer mutated and deaminated nucleic acid molecules as they appeared before fragmentation based on the pattern of mutation and deaminations in the sequence reads. For example, the first sequence reads can be compared with a reference sequence and / or sequence reads from the sample that are not mutated, to identify likely mutated bases and deaminated bases. First sequence reads can be grouped together based on patterns of mutations and deaminations contained within them. Each group can be assembled, and the assembled first sequence reads can be assigned to a positive or negative strand based on imbalance of C to T versus G to A conversions. Then, the assembled first sequence reads can be compared to a reference sequence and / or unmutated deaminated sequence reads, or unmutated not-deaminated sequence reads, in order to replace mutated bases with a base that is most likely to be accurate to the original unmutated / undeaminated sample, creating “synthetic long reads.” The synthetic long reads can then be used to improve the overall sequence assembly of the nucleic acid sample by providing long-range sequence information.

[0037] Further disclosed herein are methods of phasing sequence reads. For example, allele-specific deaminations (resulting from allele-specific methylation) can be used alongside heterozygous SNPs to phase overlapping assembled first sequence reads. Further disclosed herein are methods of determining a haplotype-specific methylation level by determining a methylation level for phased blocks of reads.Definitions

[0038] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.

[0039] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.

[0040] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.

[0041] As used herein, the term “about,” when used in reference to a measurable value such as an amount of mass, dose, time, temperature, and the like, is meant to encompass variations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.

[0042] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0043] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.

[0044] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.

[0045] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dPTP, dUTP, 5-methyl-CTP, 5-methyl- dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5- F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0046] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, areference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. One example of a publicly available reference is the human genomic sequence GRCh38 from the Genome Reference Consortium.

[0047] The term “nucleic acid sample” herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to themethods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0048] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0049] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 300 bp, and includes paired end sequence reads.

[0050] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more. A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed from multiple short sequence reads (for example, ICLRs).

[0051] As used herein, “methylation-specific deamination agent” refers to one or more agents that are used to selectively convert methylated cytosines (such as 5mC and / or 5hmC) to uracil or thymine, or alternatively to selectively convert unmethylated cytosines (uC) to uracil or thymine. Examples include engineered APOBEC deaminases, SEM-seq deaminases, TAPS reagents (including ten-eleven translocation (TET) and subsequent PCR), sodium bisulfite, and EM-seq. A methylation-specific deamination agent may broadly include a combination of agents which produce methylation-specific deamination when serially contacted with nucleic acids, for example, EM-seq where nucleic acids are first contacted with an oxidation enhancer (TET) and then with an APOBEC enzyme that is specific for nonoxidized cytosines. In some embodiments, a methylation-specific deamination agent is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% specific for either methylated cytosines or non-methylated cytosines, as is the case. In some embodiments, a methylation-specific deamination agent is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% effective for either methylated cytosines or non-methylated cytosines, as is the case. For example, by example and not by way of limitation, a methylationspecific APOBEC deaminase specific for 5mC and 5hmC may be about 96% specific for 5mC and 5hmC (only converting uC in about 4% of instances), and about 99% effective (converting about 99% of 5mC and 5hmC).

[0052] As used herein, “mutation” refers to a change in a nucleic acid sequence. For example, a mutation includes a substitution of one nucleotide base for another. A mutation may be a transition mutation (from A to G, G to A, C to T, or T to C). In another embodiment, the mutation is a transversion mutation. In some embodiments, mutations are intentionally introduced into a nucleic acid via a sample processing step which randomly introduces mutations into the nucleic acid molecule. In some embodiments, the mutations areunintentionally introduced, such as by error or unintended consequence during sample preparation or library preparation.

[0053] As used herein, a “mutated sequence read,” refers to a sequence read that contains mutations as compared to the nucleic acid template. In some embodiments, a mutated sequence read is a sequence read of a mutated nucleic acid template. In some embodiments, a mutated sequence read is obtained by sequencing a region of a mutated nucleic acid template. A mutated sequence read may have a length that is less than the length of a nucleic acid template. In some embodiments, a mutated sequence read has a length of about 100 to 600 nt, for example about 150 nt or about 300 nt. In some embodiments, mutated sequence reads may be assembled by aligning mutated sequence reads. Any method known in the art may be used for such an alignment, for example, methods as described in US20210174905A1.

[0054] As used herein, “unmutated sequence read” refers to a sequence read of the nucleic acid template. In some embodiments, an unmutated sequence read is obtained by sequencing a region of a nucleic acid template. An unmutated sequence read may have a length that is less than the length of a nucleic acid template, for example, a length of about 100 to 600 nucleotides, for example about 150 nucleotides or about 300 nucleotides.

[0055] ‘Most probable base” refers to the nucleic acid base which is determined to have the highest probability of matching the nucleic acid template (in some cases, the specific copy of a nucleic acid template) from which a mutated sequence read derives.

[0056] “Phasing” refers to a process of assigning sequence reads to a haplotype, such as based on allele-specific sequence information.Methods for determining a sequence of a nucleic acid

[0057] In an aspect, disclosed herein are methods for determining a nucleic acid sequence.

[0058] In some embodiments, the method includes initial sequencing library preparation steps. For example, the method may include fragmenting nucleic acids, and may include adding an adapter sequence to an end of the nucleic acids.

[0059] It will be readily apparent that one of ordinary skill in the art could use a variety of methods to fragment nucleic acids. For example, fragmentation methods include sonication, nebulization, hydrodynamic shearing, and tagmentation (contacting nucleic acidswith a transposase complex which fragments the nucleic acids and adds adapters to at least one end).

[0060] It will be evident to one of skill in the art that other sequencing library preparation steps may also take place before sequencing, at various points in the workflow. For example, fragmented nucleic acids may be amplified, adapters such as sequencing adapters may be added, the nucleic acids may be quantified and / or purified, as is known to those of skill in the art.

[0061] In some embodiments, fragmenting nucleic acids and adding an adapter sequence to an end of the nucleic acids comprises an initial step of tagmenting nucleic acids using a transposome complex. For example, in some embodiments, sample nucleic acids are contacted with a transposome complex comprising a transposase and an adapter nucleic acid, to form tagmented nucleic acids comprising a mosaic end sequence on at least one end of the nucleic acids. In some embodiments, during this initial step, tagmented nucleic acids are about 10 kb in length. One of skill in the art will recognize that other lengths of nucleic acids may be used at this step, for example, about 1 kb, about 5 kb, about 15 kb, about 20 kb, about 30 kb, about 50 kb, any number therebetween, or a range constructed from any of the aforementioned values. In some embodiments, the tagmentation step is a high molecular weight tagmentation.Contacting nucleic acids with a methylation-specific deamination agent

[0062] In some embodiments, the method includes contacting nucleic acids with a methylation specific deamination agent to form deaminated nucleic acids.

[0063] In some embodiments, the methylation-specific deamination agent comprises a deamination agent specific for 5-methyl cytosine (5mC) and 5- hydroxymethylcytosine (5hmC), or an unmethylated cytosine (uC) -specific deamination agent

[0064] Embodiments presented herein describe sequence reads from bisulfite converted samples. In bisulfite sequencing, DNA is chemically treated with sodium bisulfite, which results in the conversion of unmethylated cytosines to uracils, and the resulting uracils are ultimately sequenced as thymines. In contrast, the modified cytosines, 5mC and 5hmC, are resistant to bisulfite conversion, and are sequenced as cytosines. It will be understood by one of ordinary skill in the art that other methylation conversion methods can also be used to generate sequence reads for use in the methods and embodiments presented herein. Forexample, an alternative to bisulfite conversion is Enzymatic Methyl-seq (EM-seq). EM-seq and other similar methods are described in Vaisvila et al., Enzymatic methyl sequencing detects DNA methylation at single-base resolution from picograms of DNA, Genome Res. 2021 Jul; 31(7): 1280-1289, doi: 10.1101 / gr.266551.120, the content of which is incorporated herein by reference in its entirety. Briefly, EM-seq conversion uses an enzymatic method which results in the conversion of unmethylated cytosines to uracils, and the resulting uracils are ultimately sequenced as thymines. In contrast, the modified cytosines, 5mC and 5hmC, are resistant to the enzymatic conversion, and are sequenced as cytosines.

[0065] Another example of an alternative method to bisulfite conversion is TET- assisted pyridine borane sequencing (TAPS). TAPS has been described in Liu, et al. Bisulfite- free direct detection of 5 -methylcytosine and 5-hydroxymethylcytosine at base resolution. Nat Biotechnol. 2019 Apr;37(4):424-429. doi: 10.1038 / s41587-019-0041-2, and in WO 2019 / 136413, the contents of which are incorporated herein by reference in their entirety. TAPS results in conversion of 5mC and 5hmC to uracils, and the resulting uracils are ultimately sequenced as thymines. In contrast, unmethylated cytosines are resistant to the TAPS conversion, and are sequenced as cytosines. One of ordinary skill in the art will recognize that where TAPS conversion is utilized to generate sequence reads, the reference sequences described herein can be modified accordingly to reflect methylated targets where methylated cytosines are converted and sequenced as thymine (C to T conversion) and unmethylated cytosines are not converted and sequenced as cytosine. An enzymatic alternative to TAPS involves use of a modified cytidine deaminase enzyme, engineered to selectively deaminate only 5mC and 5hmC, while unmethylated cytosines are not converted. Similar to TAPS, this modified cytidine deaminase method results in conversion of 5mC and 5hmC to uracils, and the resulting uracils are ultimately sequenced as thymines. In contrast, unmethylated cytosines are resistant to the TAPS conversion, and are sequenced as cytosines.

[0066] Use of modified cytidine deaminase enzymes has been described in PCT application WO 2023 / 196572, filed April 7, 2023 and titled “Altered Cytidine Deaminases and Methods of Use”, as well as in U.S. Provisional App. Nos. 63 / 541,084, 63 / 541,097, 63 / 541,092, 63 / 541,099, and 63 / 541,077, each filed September 28, 2023, the contents of which is incorporated herein by reference in its entirety.

[0067] In some embodiments, the methylation-specific deamination comprises a deamination agent specific for methylated cytosines, such as 5mC and 5hmC. In some embodiments, the deamination agent specific for 5mC and 5hmC comprises a deaminase specific for 5mC and 5hmC. In some embodiments, the deaminase specific for 5mC and 5hmC comprises an engineered APOBEC enzyme or a SEM-seq deaminase. In some embodiments, the deamination agent specific for 5mC and 5hmC comprises TAPS reagents.

[0068] In some embodiments, the methylation-specific deamination agent comprises an uC-specific deamination agent. In some embodiments, the uC-specific deamination agent comprises sodium bisulfite, or a uC-specific deaminase. In some embodiments, the uC-specific deamination agent comprises an oxidation enhancer and an APOBEC enzyme that is specific for non-oxidized cytosines (such as is used in EM-seq).

[0069] In some embodiments, the methylation-specific deamination agent requires single-stranded nucleic acids, while the nucleic acid sample is double-stranded. For example, some deaminases may only work with ssDNA. Accordingly, in some embodiments, the method may include denaturing double-stranded nucleic acids using heat, a denaturing agent, or a combination thereof.

[0070] In other embodiments, the methylation-specific deamination agent can convert cytosines on dsDNA. For example, SEM-seq uses a cytosine deaminase which can convert 5-methylcytosine (5mC) to thymine, [doi.org / 10.1101 / 2023.06.29.547047],Introducing mutations into the deaminated nucleic acids

[0071] In some embodiments, the method includes introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids.

[0072] Embodiments relate to methods for determining a sequence of a nucleic acid template molecule using a sample processing step that intentionally introduces random mutations into a copy of the sample nucleic acid. For example, methods for introducing random mutations into a target nucleic molecule are described in U.S. Patent Application Publication No. US20210010008A1. In some embodiments, mutations are introduced by the Illumina® Complete Long Read sequencing sample preparation chemistry. The introduction of such random mutations may be useful in sequence assembly, for example, as described in U.S. Patent Application Publication No. US20210174905A1. For example, the creation of amutated copy of the nucleic acid template may be particularly useful for performing nucleic acid sequencing when a nucleic acid template includes repeat regions, for example when a nucleic acid sample includes two or more non-identical copies (for example, haplotypes) of a repeat region.

[0073] In some embodiments, the step of introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids comprises amplifying the deaminated nucleic acids with a polymerase, dNTPs, and a nucleotide analog. In some embodiments, the polymerase is a high-fidelity low-bias polymerase. In some embodiments, the mutations comprise transition mutations. In some embodiments, introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids comprises amplifying deaminated nucleic acids with a DNA polymerase such that about 1% to 15% of nucleotides in deaminated nucleic acids are mutated. In some embodiments, approximately 5% of nucleotides in the deaminated nucleic acids are mutated. In some embodiments, there are multiple rounds of amplification. In some embodiments, the DNA polymerase mutates between 0% and 3% of the nucleotides in the deaminated nucleic acids per round of replication. In some embodiments, the mutations are transition mutations. In some embodiments, introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids comprises amplifying the deaminated nucleic acids with a polymerase, dNTPs, and nucleotide analogs to randomly introduce transition mutations such that about 1% to 15% of nucleotides are mutated.

[0074] In some embodiments, the step of amplifying the nucleic acids in the presence of a nucleotide analog is followed by a step of amplifying in the absence of any nucleotide analogs (such as, with normal dNTPs). When nucleic acids that have incorporated a nucleotide analog in the previous round(s) of amplification undergo further amplification without nucleotide analogs, mutations such as transition mutations (for example, in the case of dPTP, transition mutations from C to T) are introduced in the place of the nucleotide analog.

[0075] Generally, when a low-bias DNA polymerase uses a nucleotide analog to introduce a mutation, this requires more than one round of replication. In the first round of replication the low bias DNA polymerase introduces the nucleotide analog in place of a nucleotide, and in a second round of replication, that nucleotide analog pairs with a natural nucleotide to introduce a substitution mutation in the complementary strand. The second roundof replication may be carried out in the presence of the nucleotide analog. However, the method may further comprise a step of amplifying nucleic acid molecules comprising nucleotide analogs, in the absence of any nucleotide analogs in the amplification buffer. The step of amplifying nucleic acid molecules comprising nucleotide analogs in the absence of nucleotide analogs in the amplification buffer may be carried out using a low-bias DNA polymerase.

[0076] Accordingly, some embodiments of systems and methods disclosed herein include steps of obtaining one or more nucleic acid samples (including obtaining one nucleic acid sample that is split into two nucleic acid samples and including two nucleic acid samples from the same subject). Some embodiments include randomly introducing mutations into nucleic acid molecules in one or more nucleic acid samples and producing a mutated nucleic acid sample. Some embodiments include keeping a second nucleic acid sample unmutated by not randomly introducing mutations into nucleic acid molecules in the nucleic acid sample.Amplifying and fragmenting the mutated and deaminated nucleic acids

[0077] In some embodiments, the method includes amplifying and fragmenting the mutated and deaminated nucleic acids.

[0078] A step of amplifying the mutated and deaminated nucleic acids could be carried out using any appropriate method, for example, PCR. In some embodiments, amplifying comprises a bottlenecking (suppression) PCR as further described below.

[0079] The step of fragmenting mutated and deaminated nucleic acids could be carried out using any appropriate method. For example, fragmentation can be carried out using restriction digestion or using PCR with primers complementary to at least one internal region of the at least one mutated target nucleic acid molecule. In some embodiments, fragmentation is carried out using a technique that is not sequence-specific, for example, sonication or nebulization. In some embodiments, fragmentation is carried out by tagmentation, as further described below.

[0080] Additional library preparation steps may also be performed before sequencing to generate fragments of a suitable length for a desired sequencing technology. For example, WO 2023 / 230552, hereby incorporated by reference in its entirety, discloses methods for preparing a library of mutated nucleic acids. In some embodiments, the method includes normalizing the library of mutated and deaminated nucleic acids (library normalization). Thisstep may be accomplished using methods known to those of skill in the art, including using standard workflows for libraries, for example, bead-based normalization, enzymatic normalization, heat-denaturation and re-annealing (Cot-based) normalization, quantificationbased normalization, or normalization by hybridization.

[0081] In some embodiments, the method includes the amplifying the mutated and deaminated nucleic acids in a bottlenecking (suppression) PCR step. For example, in this step, a defined quantity of the purified mutagenesis product may be amplified to create many copies of each unique template. The amount of starting material in the bottlenecking PCR may determine the number of long templates available for sequencing, and may be controlled through dilution of the mutagenesis sample. In some embodiments, suppression (“bottlenecking” or “bottleneck”) PCR acts on larger fragments. In some embodiments, suppression PCR entails appending complementary sequences on the 5’ and 3’ ends of the same DNA molecule, such that during a PCR annealing step, there is a direct competition between annealing of a primer and annealing of opposite ends of the same DNA fragment. When the PCR primer anneals, extension proceeds as normal, and the fragment is amplified. When opposite ends anneal, for example by forming a hairpin, there is no templated 3' hydroxyl to extend, and amplification does not occur. In some embodiments, for shorter fragments, the opposite ends of the same fragment are closer together and therefore more likely to find each other and anneal. Under optimized conditions, this can lead to preferential amplification of longer fragments. Aspects of suppression PCR useful with embodiments provided herein are described in Dai, Z-M, et al (2006) I. of Biotech 128:435-443; and Rand K.N. et al., (2005) N.A. Res.33:el27 which are incorporated by reference in their entireties.

[0082] In some embodiments, the mutated and deaminated nucleic acids are contacted with transposome complexes in a tagmentation step. The tagmentation may be second tagmentation step, for example, as described above there may be initial sequencing library preparation steps. In some embodiments, the tagmentation is a low molecular weight tagmentation step. For example, in some embodiments, low molecular weight tagmentation is performed using Nextera DNA Flex (“Illumina DNA Prep”; Illumina Inc., San Diego, CA) according to manufacturer instructions.

[0083] In some embodiments, the method includes adding an index sequence to an end of the mutated and deaminated nucleic acids. Index sequences may be added by manymethods known to those of skill in the art, such as amplification with primers that include an index sequence, ligation, or tagmentation. In some embodiments, the method includes adding an index during a PCR reaction. For example, in some embodiments, tagmentation products are then amplified using primers that contain sample-specific index sequences, according to manufacturer instructions (“Illumina DNA Prep”; Illumina Inc., San Diego, CA).Sequencing the mutated and deaminated nucleic acids

[0084] In some embodiments, the method includes sequencing the mutated and deaminated nucleic acids to obtain first sequence reads. It will be readily apparent that one of ordinary skill in the art could use a variety of methods to sequence the mutated and deaminated nucleic acids. Sequencing may be accomplished by any method known to those of skill in the art, including next-generation sequencing (NGS). In some embodiments, the first sequence reads are short sequence reads, for example, are less than about 500 bp in length. In some embodiments, the first sequence reads are between 100 and 400 bp in length. In some embodiments, the first sequence reads are paired-end sequence reads.

[0085] Some embodiments herein also include steps of sequencing deaminated unmutated nucleic acids, such as nucleic acids which have been contacted with a methylationspecific deamination agent but that have not had mutations introduced into them. Some embodiments herein also include steps of sequencing unmutated nucleic acids that are not deaminated, such as nucleic acids which have not been contacted with a methylation-specific deamination agent and that have not had mutations introduced into them. Sequencing of deaminated and mutated nucleic acids, and unmutated nucleic acids, may be performed simultaneously or sequentially.Assembling the first sequence reads

[0086] In some embodiments, the method includes assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

[0087] FIG. 1A is a flow diagram that schematically illustrates an exemplary method 100 for assembling sequence reads based on patterns of mutations and deaminations. In some embodiments, the method 100 is implemented on a computer. The method 100 maybe embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. When the method 100 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device.

[0088] As shown in FIG. 1 A, the method 100 for assembling sequence reads based on patterns of mutations and deaminations may start from start block 110. The method 100 may proceed to block 120, wherein first sequence reads from mutated and deaminated nucleic acids are received, for example, retrieved from a storage. The first sequence reads may be generated from mutated and deaminated nucleic acids, as further described above. The sequence reads may be in any digital file format.Comparing Sequence Reads with a Reference to Identify Mutated Bases

[0089] The method 100 may proceed to block 130, wherein first sequence reads are compared with a reference sequence and / or with second sequence reads to identify likely mutated bases and / or deaminated bases. In some embodiments, the second sequence reads comprise sequence reads from deaminated unmutated nucleic acids, or from unmutated nucleic acids that are not deaminated.

[0090] For example, at this step, first sequence reads may be aligned to a reference sequence. In some embodiments, to identify mutated bases corresponding to the mutations introduced into the deaminated nucleic acids, first sequence reads are compared with a reference sequence. For example, the first sequence reads may be aligned to a reference sequence such as a hg38 reference sequence (available at ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / ). For example, if the mutation process introduces transition mutations, the comparison may identify transition differences between the first sequence reads and the reference sequence and these differences may be flagged.

[0091] In some embodiments, first sequence reads are compared with unmutated second sequence reads. In some embodiments, the second sequence reads include sequence reads from deaminated unmutated nucleic acids. For example, after a step of contacting nucleic acids with a deamination agent to form deaminated nucleic acids (further described above), a portion of the deaminated nucleic acids may be split off, and may proceed to sequencingwithout having mutations introduced — thus providing deaminated unmutated nucleic acids and sequence reads thereof. For example, the comparison may identify unmatched bases between the first sequence reads and second sequence reads, and these unmatched bases may be flagged.

[0092] In some embodiments, the second sequence reads includes sequence reads from unmutated nucleic acids that are not deaminated, for example nucleic acids from the same sample which have not been deaminated and have not have mutations introduced into them. These unmutated, not deaminated nucleic acids may proceed to sequencing to obtain sequence reads from unmutated nucleic acids that are not deaminated.

[0093] It will be readily apparent that one of ordinary skill in the art could use a variety of methods to identify likely mutated bases and / or deaminated bases based on comparison with a reference sequence and / or second sequence reads. For example, in some embodiments, a probability that a base corresponds to a mutation or deamination may be calculated for each base in the first sequence reads. In some embodiments, bases in the first sequence reads may be flagged or stored as likely mutation / deaminated bases based on the comparison with a reference sequence and / or second sequence reads, for example if the probability is above a predetermined threshold.Grouping First Sequence Reads

[0094] The method 100 may then proceed to block 140, wherein first sequence reads are grouped based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids.

[0095] Accordingly, in some embodiments, assembling the first sequence reads includes grouping first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids, to generate groups of first sequence reads.

[0096] It will be readily apparent that one of ordinary skill in the art could use a variety of methods to form groups of first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids. For example, the groups may be groups of sequence reads which are determined to be likely (for example, to have a probability over a predetermined threshold) to have originated from the same original mutated and deaminated nucleic acids.

[0097] For example, in some embodiments, the first sequence reads are analyzed, for example compared to second sequence reads which are unmutated (and optionally may or may not be deaminated), to determine k-mers within the first sequence reads that correspond to mutation patterns. Information related whether a base likely corresponds to a mutation or deamination (see above) may also be used to determine mutation patterns used to group first sequence reads.

[0098] In some embodiments, first sequence reads and second sequence reads are compared and analyzed to build a weighted network of first sequence read pairs based on the likely mutation sites and deamination sites. In some embodiments, C to T and G to A mutations / deaminations are downweighted based on CpG context (likely methylation site), occurrence across reads, or strand information if available.

[0099] First sequence reads may be grouped together on the basis of including a common k-mer that corresponds to a mutation pattern. In some embodiments, Louvain and / or Markov clustering is used to group first sequence reads. W02020035669A1, which is hereby incorporated by reference in its entirety, describes methods for pre-clustering sequence reads. Furthermore, WO2021064365A1, which is hereby incorporated by reference in its entirety, describes methods of determining a measure correlated to the probability that two mutated sequence reads derive from the same sequence comprising mutations.Assembling Each Group of First Sequence Reads

[0100] The method 100 may then proceed to block 150, wherein each group of first sequence reads is assembled.

[0101] It will be readily apparent that one of ordinary skill in the art could use a variety of methods to assemble each group of first sequence reads. For example, in some embodiments, the assembly is guided by reference sequence. In some embodiments, the assembly is guided by nonmutated second sequence reads. In some embodiments the assembly is not based on a reference.

[0102] In some embodiments, first sequence reads are aligned to a reference sequence, for example, hg38. In some embodiments, alignments of nearby reference regions may be merged. In some embodiments, the alignment of first sequence reads to the reference sequence may guide assembly of each group of first sequence reads. Exemplary methods forassembly of mutated sequence reads are described in W02020035669A1, which is hereby incorporated by reference in its entirety.Assigning Assembled Sequence Reads to a Strand

[0103] The method 100 may then proceed to block 160, wherein assembled first sequence reads are assigned to a strand (such as one of two strands of a double-stranded nucleic acid) based on nucleotide composition. For example, nucleotide composition of assembled first sequence reads may be analyzed to determine a probability that the assembled first sequence read corresponds to a Watson strand or a Crick strand.

[0104] In some embodiments, assembled first sequence reads are assigned to a strand based on cytosine (C) to thymine (T) and guanine (G) to adenine (A) imbalance arising from deamination of 5mC to T or deamination of uC to T. For example, in some embodiments, assembled first sequence reads with higher C to T conversions and lower G to A conversions are assigned to the Watson strand (positive or + strand), while assembled first sequence reads with lower C to T conversions and higher G to A conversions are assigned to the Crick strand (negative or - strand). In some embodiments, assembled first sequence reads that do not fit into either of the two categories above (for example, do not have an imbalance above a threshold) may be discarded or marked as an unknown strand of origin.

[0105] In some embodiments, as described above, assigning to a strand may take place after assembling the first sequence reads. In some embodiments, a step of assigning to a strand may take place earlier in the workflow, such as before grouping first sequence reads or before assembling first sequence reads. These embodiments may include embodiments where nucleic acids are contacted with an unmethylated cytosine (uC) -specific deamination agent, such as via EM-seq or bisulfite seq, because a greater number of C to T conversions may take place and allow for strand inference based on the first sequence reads themselves. Accordingly, in some embodiments, the method includes assigning first sequence reads or groups of first sequence reads to a strand based on nucleotide composition.Replacing mutated bases

[0106] The method 100 may then proceed to block 170, wherein mutated bases are replaced. For example, bases in the assembled first sequence reads, which were identified asbeing likely to correspond to a mutation or deamination in the mutated and deaminated nucleic acids, may be replaced with a base from a reference sequence, bases from deaminated unmutated sequence reads, or bases from sequence reads from unmutated nucleic acids that are not deaminated.

[0107] In some embodiments, mutated bases are replaced with non-mutated bases by comparing assembled first sequence reads with a reference sequence, deaminated unmutated sequence reads, and / or sequence reads from unmutated nucleic acids that are not deaminated. In some embodiments, the non-mutated base is a most-probable base.

[0108] In some embodiments, C to T transitions on the positive strand sequence and G to A transitions on the negative strand sequence are not processed for replacement with a non-mutated base, for example when non-mutated not-deaminated sequence reads are used as second sequence reads, in order to preserve methylation information as these may be likely deaminated bases. In such embodiments, unmutated, not-deaminated sequence reads may be used as second sequence reads for analysis and comparison with the assembled first sequence reads.

[0109] In some embodiments, all likely mutation and / or deaminated bases are processed for replacement with a non-mutated base, and C to T transitions on the positive strand sequence or to G to A transitions on the negative strand sequence are not excluded (for example, all likely mutation and / or deaminated bases may be processed). In such embodiments, deaminated unmutated sequence reads may be used as second sequence reads for analysis and comparison with the assembled first sequence reads.

[0110] For example, in some embodiments, the process for replacing mutated bases may proceed as follows. First, assembled first sequence reads are aligned to a reference sequence and / or unmutated deaminated second sequence reads. In some embodiments, a most- probable base is statistically determined at the position of each likely mutated base or deaminated base in the assembled first sequence reads by analysis of nonmutated second sequence reads, wherein the most-probable base is a nucleotide that has the highest determined probability of corresponding to the nucleic acids before they are deaminated and / or before mutations are introduced. In some embodiments, each likely mutation or deaminated base in the assembled first sequence reads is replaced with the most-probable base.[oni] Methods for replacing mutated bases with non-mutated bases are described in Int. App. No. PCT / US2024 / 018539, which is hereby incorporated by reference in its entirety. Similar techniques may be used to replace mutated bases at block 170.Methods of identifying a sequence variant

[0112] Further described herein are methods of identifying a sequence variant. In some embodiments, the method includes determining a nucleic acid sequence as described herein, and identifying a sequence variant in the nucleic acid sequence by comparing to a reference sequence.

[0113] In some embodiments, the sequence variant is identified as being associated with a specific strand of a double-stranded nucleic acid. For example, as described further above, a first sequence read, a group of first sequence reads, and / or assembled first sequence reads may be assigned to a strand based on cytosine (C) to thymine (T) and guanine (G) to adenine (A) imbalance arising from deamination of 5mC to T. As further described herein, the assignment may be based on a statistical inference based on nucleotide imbalance.

[0114] Methods for variant calling are described in PCT7US2024 / 035562 filed June 26, 2024, which is hereby incorporated by reference in its entirety. For example, the method can include identifying, for a target genomic sample, nucleotide reads comprising one or more nucleobases converted by a methylation sequencing assay; determining an estimated methylation-level value for a cytosine base at a genomic coordinate based on prior genotype probabilities for the target genomic sample at the genomic coordinate and observed nucleobases at the genomic coordinate within the nucleotide reads; generating, utilizing a variant call model, posterior genotype probabilities for the target genomic sample at the genomic coordinate based on the estimated methylation-level value and base-call-quality metrics for the observed nucleobases; and generating, based on the posterior genotype probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate.Methods of phasing and determining a 5mC or uC haplotype

[0115] Further described herein are methods of phasing sequence reads and methods of determining a 5mC haplotype or a uC haplotype.

[0116] Read-based phasing algorithms use informative sites (such as singlenucleotide polymorphisms) contained within sequence reads to connect sequence reads together into haplotypes. Allele-specific methylation sites which are deaminated as described herein provide additional haplotype-informative sites that increase phasing capabilities and provide for the ability to determine the phase of sequence reads in larger blocks. The ability to assemble first sequence reads based on introduced mutations also improves the ability to phase sequence reads into longer haplotype blocks.

[0117] In some embodiments, the methods of phasing include receiving first sequence reads from mutated and deaminated nucleic acids, and phasing the first sequence reads using 5-base phasing with 5mC / 5hmC (or uC, when uC has been specifically deaminated instead of 5mC / 5hmC) as a fifth base. For example, the method may include connecting overlapping first sequence reads or assembled first sequence reads to each other based on heterozygous single-nucleotide polymorphisms (SNPs) and allele-specific deaminations, thereby determining phased blocks of first sequence reads or assembled first sequence reads.

[0118] For example, in some embodiments, a phasing method uses differences in the distribution of methylation and bases between haplotypes.

[0119] Methylation positions suitable to use for phasing cannot be determined in the same way as sequence variants. Sequence variants suitable for phasing are heterozygote variants and can be easily identified from an allele frequency of approximately 0.5 as present on one haplotype, not on the other. But the methylation state at a genome position on a haplotype is much less likely to be consistent between reads originating from different DNA molecules of the same haplotype (for example from different cells). A methylation frequency of 0.5 can be caused by a methylation frequency of 0.5 in both haplotypes and not be a suitable site to use for phasing. Thus, in some embodiments, a phasing method includes a method to identify informative allele-specific methylation sites.

[0120] An example phasing algorithm using methylation has the following steps: step i) phase all reads based on sequence positions, step ii) identify usable methylation sites based on haplotype-specific methylation at positions spanned by phased reads, step iii) rephase any previously unphased reads using sequence positions and newly identified allele specific methylation sites, and step iv) iterate until no additional allele-specific methylation sites are identified.

[0121] It is not clear that there are only 2 methylation haplotypes even in a diploid organism. For example, different cell types are known to have different methylation patterns. Thus, in some embodiments, the methylation phasing method allows that greater than 2 haplotypes may have generated the observed read methylation. In some embodiments, the phasing method estimates the number of methylation haplotypes, counts the number of cell types in the sample, and estimates cell type proportions.

[0122] Phasing methods may take into account methylation on a per-base level and / or in subsections of the nucleic acid sequence that include multiple bases. For example, methylation per base may be noisier than the DNA sequence, for example because there are multiple methylation states in different cell types which may be included in the sample. By considering the methylation state of multiple bases, a more robust allele or cell specific methylation signal may be achieved for use in phasing.

[0123] The method may also include determining a nucleic acid sequence as described herein.

[0124] In some embodiments, disclosed herein are methods of determining a haplotype-specific methylation level. In some embodiments, the method includes phasing first sequence reads as described herein; and determining a haplotype-specific methylation level. For example, once first sequence reads have been phased into blocks (which correspond to haplotypes), a methylation level (such as levels of 5mC and / or 5hmC) may be determined for each phased block.

[0125] In some embodiments, the method further includes storing one or more haplotype-specific methylation levels in a digital file. In some embodiments, the digital file is a VCF file. In some embodiments, methylation levels per haplotype are encoded in the phase set (PS) field of the VCF file.Kits for determining a sequence of a nucleic acid

[0126] Further described herein are kits for determining a sequence of a nucleic acid. In some embodiments, the kit includes a methylation-specific deamination agent, and one or more mutagenesis agents.

[0127] The methylation-specific deamination agent may be any of the methylationspecific deamination agents disclosed herein. For example, in some embodiments, themethylation-specific deamination agent is specific for 5mC and 5hmC, such as an engineered APOBEC enzyme, a SEM-seq deaminase, or TAPS reagents. In some embodiments, the methylation-specific deamination agent is specific for uC, such as sodium bisulfite, a uC- specific deaminase, or an oxidation enhancer and an APOBEC enzyme that is specific for nonoxidized cytosines.

[0128] The one or more mutagenesis agents may be any reagents which can introduce mutations into nucleic acids, such as is described herein. For example, the one or more mutagenesis agents may include a DNA polymerase, dNTPs, and a nucleotide analog. In some embodiments, the nucleotide analog is dPTP.Systems and Computer-Readable Media

[0129] Further disclosed herein are electronic systems for assembling sequence reads based on patterns of mutations and deaminations. In some embodiments, the system includes a processor configured to perform a method comprising: receiving first sequence reads from mutated and deaminated nucleic acids; comparing first sequence reads with a reference sequence and / or with reference sequence reads to identify likely mutated bases and / or deaminated bases; grouping first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids; assembling each group of first sequence reads; assigning assembled first sequence reads to a strand based on nucleotide composition; and replacing mutated bases.

[0130] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive first sequence reads from mutated and deaminated nucleic acids; compare first sequence reads with a reference sequence and / or with reference sequence reads to identify likely mutated bases and / or deaminated bases; group first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids; assemble each group of first sequence reads; assign assembled first sequence reads to a strand based on nucleotide composition; and replace mutated bases

[0131] FIG. 2A illustrates a diagram of an environment in which a system for determining a sequence of a nucleic acid can operate in accordance with one or moreimplementations. The following paragraphs describe the system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 2A illustrates a schematic diagram of a computing system 2000 in which a deamination / mutation sequencing application 2106 operates in accordance with one or more implementations. As illustrated, the computing system 2000 includes one or more server device(s) 2102 connected to a user client device 2108, a local device 2118, and a sequencing device 2114 via a network 2112. The network 2112 can comprise any suitable network over which computing devices can communicate.

[0132] As shown in FIG. 2A, the computing system 2000 includes the server device(s) 2102. In various implementations, the server device(s) 2102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 2102 receive various data from the sequencing device 2114, such as data from a sample genome and / or sequence reads. The server device(s) 2102 may also communicate with the user client device 2108. In particular, the server device(s) 2102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 2108.

[0133] As shown, the server device(s) 2102 includes a sequencing application 2110. In general, the sequencing application 2110 analyzes the data (such as call data) received from the sequencing device 2114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 2110 can receive raw data from the sequencing device 2114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 2110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0134] As also shown, the sequencing application 2110 includes the deamination / mutation sequencing application 2106. As described below, in some embodiments, the deamination / mutation sequencing application 2106 can determine a sequence of a nucleic acid. For example, in some embodiments, the deamination / mutation sequencing application 2106 receives first sequence reads from mutated and deaminated nucleic acids; and assembles the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

[0135] While the sequencing application 2110 has been described as including the deamination / mutation sequencing application 2106, other systems or methods may be included within the sequencing application 2110, such as an application to determine single nucleotide polymorphisms or to determine a 5mC or uC haplotype (not illustrated).

[0136] Moreover, while the deamination / mutation sequencing application 2106 is described being implemented on the server device(s) 2102, as part of the sequencing application 2110, in some implementations, the deamination / mutation sequencing application 2106 is implemented by (such as located entirely or in part) on the user client device 2108, the sequencing device 2114, and / or the local device 2118. As mentioned, in some implementations, deamination / mutation sequencing application 2106 is implemented by one or more other components of the computing system 2000, such as the sequencing device 2114. In particular, the deamination / mutation sequencing application 2106 can be implemented in a variety of different ways across the server device(s) 2102, the network 2112, the user client device 2108, the local device 2118, and the sequencing device 2114.

[0137] As further shown in FIG. 2A, the computing system 2000 includes the user client device 2108. In various implementations, the user client device 2108 can generate, store, receive, and send digital data. In particular, the user client device 2108 can receive the data from the sequencing device 2114. As further illustrated, the user client device 2108 includes a sequencing application 2110. The sequencing application 2110 may be a web application or a native application stored and executed on the user client device 2108 (for example, a mobile application, desktop application, or web application). The sequencing application 2110 can receive data from the sequencing application 2110 and / or deamination / mutation sequencing application 2106. For example, the user client device 2108 can receive variant call files and / or alignment files from the sequencing application 2110.

[0138] The sequencing application 2110 can also include instructions that (when executed) cause the user client device 2108 to receive data from the deamination / mutation sequencing application 2106 and present data from the sequencing device 2114 and / or the server device(s) 2102. Furthermore, the sequencing application 2110 can instruct the user client device 2108 to display data for variant calls, such as nucleobase calls or an indication of a methylation level. Indeed, the user client device 2108 can display nucleobase call results for a genome sample and / or an indication of a level of 5mC and / or uC for a haplotype.

[0139] As further shown in FIG. 2A, the computing system 2000 includes the sequencing device 2114. In various implementations, the sequencing device 2114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 2114 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 2114. More particularly, the sequencing device 2114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 2114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 2112, in some implementations, the sequencing device 2114 bypasses the network 2112 and communicates directly with the user client device 2108.

[0140] As further depicted in FIG. 2A, in some implementations, the server device(s) 2102 includes a distributed collection of servers, where the server device(s) 2102 include several server devices distributed across the network 2112 and located in the same or different physical locations. For instance, the server device(s) 2102 can be implemented, in whole or in part, on the local device 2118. To illustrate, the local device 2118 may implement the sequencing application 2110 and / or the deamination / mutation sequencing application 2106. Further, the server device(s) 2102 and / or the local device 2118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0141] The user client device 2108 illustrated in FIG. 2 A can include various types of client devices. For example, in some implementations, the user client device 2108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 2108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0142] Though FIG. 2A illustrates the components of the computing system 2000 communicating via the network 2112, in certain implementations, the components of computing system 2000 can also communicate directly with each other, bypassing the network 2112. For instance, in some implementations, the user client device 2108 communicates directly with the sequencing device 2114. Additionally, in some implementations, the user client device 2108 communicates directly with the deamination / mutation sequencingapplication 2106 and / or the server device(s) 2102. In some implementations, the user client device 2108 communicates directly with the local device 2118. Moreover, the deamination / mutation sequencing application 2106 can access one or more databases housed on or accessed by the server device(s) 2102 or elsewhere in the computing system 2000.

[0143] FIG. 2B is a block diagram of an exemplary server device 2102 that may be used in connection with the computing system 2000 of FIG. 2A. The server device 2102 may be configured to determine a sequence of a nucleic acid. The general architecture of the server device 2102 depicted in FIG. 2B includes an arrangement of computer hardware and software components. The server device 2102 may include many more (or fewer) elements than those shown in FIG. 2B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 2102 includes a processing unit 210, a network interface 220, a computer readable medium drive 230, an input / output device interface 240, a display 250, and an input device 260, all of which may communicate with one another by way of a communication bus. The network interface 220 may provide connectivity to one or more networks or computing systems. The processing unit 210 may thus receive information and instructions from other computing systems or services via a network. The processing unit 210 may also communicate to and from memory 270 and further provide output information for an optional display 250 via the input / output device interface 240. The input / output device interface 240 may also accept input from the optional input device 260, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0144] The memory 270 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 210 executes in order to implement one or more embodiments. The memory 270 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 270 may store an operating system 272 that provides computer program instructions for use by the processing unit 210 in the general administration and operation of the server device 2102. The memory 270 may store a reference genome 273, such as for use by the sequencing application 2110. The memory 270 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0145] For example, in one embodiment, the memory 270 includes a sequencing application 2110, which may include a deamination / mutation sequencing application 2106. The deamination / mutation sequencing application 2106 can perform the methods disclosed herein. In addition, memory 270 may include or communicate with the data store 290 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference genomes.

[0146] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0147] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0148] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0149] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0150] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0151] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument.In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0152] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0153] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0154] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layoutsuch as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0155] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.EXAMPLES

[0156] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1

[0157] In the following example, a nucleic acid sequence, phasing, and haplotypespecific methylation level is determined for sample DNA.

[0158] An exemplary workflow is shown in FIG. 3. As shown in FIG. 3, sample DNA is obtained and is tagmented to produce fragments of about 10 kb in length with mosaicend sequencing adapters. The sample is purified in a post-tagmentation clean up step. The tagmented fragments are denatured and treated with an engineered APOBEC deaminase that is specific for 5mC and 5hmC. The deaminated DNA is split into two samples. The first sample is amplified with a high-fidelity, low-bias PrimeSTAR® polymerase in the presence of the nucleotide analog dPTP for 6 cycles and in the absence of dPTP for 15 cycles to introducetransition mutations in about 5% of nucleotides. The mutagenized nucleic acids are purified and put through a library normalization step. The mutagenized nucleic acids are amplified in a bottlenecking PCR step. Then, to fragment the mutagenized nucleic acids, low molecular weight tagmentation is performed using Nextera DNA Flex (“Illumina DNA Prep”; Illumina Inc., San Diego, CA) according to manufacturer instructions. Tagmentation products are then amplified using primers that contain sample-specific index sequences, according to manufacturer instructions (“Illumina DNA Prep”; Illumina Inc., San Diego, CA). A sequencing library is also prepared for the second sample, which does not undergo mutagenesis. Each sample undergoes sequencing on an Illumina® NGS sequencing system, producing first sequence reads and second sequence reads.

[0159] The first sequence reads are aligned to reference genome hg38 and transition differences are flagged. The first sequence reads are also compared with the second sequence reads and unmatched bases are flagged. The transition differences and unmatched bases are analyzed to identify likely mutation and deamination sites in the first sequence reads.

[0160] First sequence reads and second sequence reads are compared and analyzed to build a weighted network of first sequence read pairs based on the likely mutation sites and deamination sites. Further analysis, including Louvain-like community detection and Markov clustering, creates groups of first sequence reads that are statistically likely to have originated from the same template mutated and deaminated nucleic acid.

[0161] First sequence reads are aligned to hg38, and each group of first sequence reads is assembled using a reference-guided assembly. Nucleotide composition is analyzed for the assembled first sequence reads, and assembled first sequence reads are assigned to the positive (+) strand if they have a higher C to T than G to A count, while assembled first sequence reads are assigned to the negative (-) strand if they have a higher G to A than C to T count.

[0162] Assembled first sequence reads are aligned to the reference genome hg38 and to second sequence reads. A minimizer sequence index is generated based on the first sequence reads and second sequence reads and the alignment of the assembled first sequence reads to the second sequence reads is updated. A Bayesian most-probable base is determined for each of the likely mutation sites and deamination sites within the assembled first sequencereads based on analysis of the second sequence reads, thereby determining synthetic long reads. A VCF fde is created which includes the sequence of the sample DNA.

[0163] Synthetic long reads are phased using allele-specific deamination sites by connecting overlapping synthetic long reads to each other based on single-nucleotide polymorphisms and allele-specific deaminations, thereby determining phased blocks of synthetic long reads.

[0164] A methylation level is determined for each haplotype. The haplotypespecific methylation level is encoded in the phase set (PS) field of the VCF file.Example 2

[0165] In the following example, two options for second sequence reads are compared in the context of the step of replacing mutations. The previous steps of the method are as described in Example 1.

[0166] As shown in FIG. 4, in option 1, assembled first sequence reads are compared to unmutated, not-deaminated short reads (of less than 500 bp) as second sequence reads. Base 401 is a C to T methylation conversion (deamination). Base 402 represents an unmethylated C to T transition that happened as a result of the step of introducing mutations. Computational processing as described herein is used to replace mutated bases, but C to T transitions on the positive strand (base 401, base 402) and G to A transitions on the negative strand are not processed.

[0167] In Option 2, assembled first sequence reads are compared to deaminated unmutated short reads as second sequence reads. As above, base 401 is a C to T methylation conversion (deamination) and base 402 represents an unmethylated C to T transition that happened as a result of the step of introducing mutations. Base 403 is a methylated cytosine that was not converted by the methylation-specific deaminase due to enzyme error in the mutated, deaminated sample but was converted in the unmutated, deaminated sample ref reads. Computational processing as described herein is used to process all likely mutation and deaminated bases for potential replacement, based on analysis of and comparison with the unmutated deaminated ref reads.

[0168] FIG. 4 illustrates the tradeoff of these two approaches. With the first option, all methyl-seq C to T conversions (deaminations) are preserved in the synthetic long read fordownstream processing such as methylation level analysis. However, because some C to T mutations are also preserved, there is an approximate 1.25% increase in the false positive rate where cytosine was mutated to thymine during the mutagenesis step (based on a 5% overall mutation rate where all bases are mutated at approximately equal rates).

[0169] With the second option, C to T mutations from the mutagenesis are removed. The methyl-seq C to T conversions (deaminations, base 401) are preserved in the synthetic long read for downstream processing. However, because of partial effectivity of the methylation-specific deaminase or because the position may be partially methylated biologically (such as in different cell types included in the sample or in different haplotypes), there is a drop in the quality score at partially methylated positions (base 402, base 403 which were deaminated in some nucleic acid molecules but not others), and there is no increase in false positive methylation rate.Example 3

[0170] In the following example, strand inference accuracy was tested on simulated sequence reads of various lengths, with methylated cytosines converted to thymines. If a simulated sequence read had a higher C to T than G to A count, it was assigned to the positive (+) strand. Inversely, if a simulated sequence read had a higher G to A than C to T count, it was assigned to the negative (-) strand.

[0171] In the simulated data set, it was known which strand the sequence reads were generated from. As shown in FIG. 5, the accuracy of strand assignment improved with the length of the simulated sequence read. The data presented in FIG. 5 suggests that a sequence read (such as a synthetic long read) of at least about 3 kb can accurately be assigned to a read strand.Example 4

[0172] In the following example, assembled first sequence reads of various sizes were phased using heterozygous SNPs and using various levels of allele-specific methylation, using 5mC as a fifth base for phasing purposes.

[0173] In this example, the experimentally-derived genome positions of allelespecific sequence and allele-specific methylation were used, but reads were not explicitly simulated.

[0174] In this example, allele-specific methylation is defined based on the “methylation imbalance” at each genome position between reads derived from each haplotype. The “methylation imbalance” at a genome position is the difference in the proportion of methylated reads between the more-methylated haplotype and less-methylated haplotype at that position.

[0175] Different thresholds for allele-specific methylation were used to classify positions as having allele-specific methylation or not (the threshold used is the x-axis of FIG. 6). If the difference in methylation imbalance at a genome position between haplotypes was greater than the threshold value, then that position was considered to have allele-specific methylation.

[0176] It was assumed that if allele-specific methylation exists at a position, it could be used for phasing as if it were a sequence difference between the haplotypes.

[0177] To simulate the phase blocks recoverable from the genome positions where differences were found between haplotypes, differentiating positions (allele-specific methylation and allele-specific sequence) within “fragment size” (legend of FIG. 6) distance of each other on the genome are joined in the same phase block. This reflects the assumption that they could be spanned by a read or read pair given sufficient read coverage, and so can be phased relative to each other.

[0178] FIG. 6 shows that 5-base phasing using simulated assembled first sequence reads with mutations replaced (synthetic long reads) of 5,000 or 10,000 bp produced much larger phased blocks than blocks which would be produced by 4-base phasing using conventional long-read sequencing.Other Considerations

[0179] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, suchconditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0180] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0181] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (such as a measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.

[0182] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.

[0183] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope. bi

[0184] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

[0185] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.

Claims

WHAT IS CLAIMED IS:

1. A method for determining a sequence of a nucleic acid, the method comprising: contacting nucleic acids with a methylation-specific deamination agent, to form deaminated nucleic acids; introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids; amplifying and fragmenting the mutated and deaminated nucleic acids; sequencing the mutated and deaminated nucleic acids to obtain first sequence reads; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

2. The method of claim 1, wherein the methylation-specific deamination agent comprises a deamination agent specific for 5-methyl cytosine (5mC) and 5- hydroxymethylcytosine (5hmC), or an unmethylated cytosine (uC) -specific deamination agent.

3. The method of claim 1 or claim 2, wherein assembling the first sequence reads comprises comparing the first sequence reads with a reference sequence or with second sequence reads to identify likely mutated bases or deaminated bases.

4. The method of claim 3, wherein the second sequence reads comprise sequence reads from deaminated unmutated nucleic acids, or from unmutated nucleic acids that are not deaminated.

5. The method of any of claims 1-4, wherein assembling the first sequence reads comprises: grouping first sequence reads based on patterns of mutations and deaminations introduced into the mutated and deaminated nucleic acids, to generate groups of first sequence reads; and assembling each group of first sequence reads.

6. The method of claim 5, wherein the method further comprises aligning assembled first sequence reads to a reference sequence or to second sequence reads, whereinthe second sequence reads comprise sequence reads from deaminated unmutated nucleic acids, or from unmutated nucleic acids that are not deaminated.

7. The method of any of claims 1-6, wherein the method comprises assigning first sequence reads, groups of first sequence reads, or assembled first sequence reads to a strand based on nucleotide composition.

8. The method of claim 7, wherein the method comprises assigning assembled first sequence reads to a strand based on cytosine (C) to thymine (T) and guanine (G) to adenine (A) imbalance arising from deamination of 5mC and 5hmC to T or deamination of uC to T.

9. The method of any of claims 1-8, wherein the method further comprises replacing a likely mutated base with a base from a reference sequence, a base from deaminated unmutated sequence reads, or from unmutated nucleic acids that are not deaminated.

10. The method of any of claims 1-9, wherein the method further comprises identifying a sequence variant in the nucleic acid sequence.

11. The method of claim 10, wherein the sequence variant is identified as being associated with a specific strand of a double-stranded nucleic acid.

12. The method of any of claims 1-11, wherein the method further comprises phasing the first sequence reads.

13. The method of any of claims 1-12, wherein the method further comprises phasing the first sequence reads using 5-base phasing with 5mC and 5hmC, or uC, as a fifth base.

14. A method of determining a haplotype-specific methylation level, comprising: phasing first sequence reads according to the method of claim 12 or claim 13; and determining a haplotype-specific methylation level.

15. The method of claim 2, wherein the deamination agent specific for 5mC and 5hmC comprises a deaminase specific for 5mC and 5hmC.

16. The method of claim 0, wherein the deaminase specific for 5mC and 5hmC comprises an engineered APOBEC enzyme or a SEM-seq deaminase.

17. The method of claim 2, wherein the deamination agent specific for 5mC and 5hmC comprises TAPS reagents.

18. The method of claim 2, wherein the uC-specific deamination agent comprises sodium bisulfite, or a uC-specific deaminase.

19. The method of claim 2, wherein the uC-specific deamination agent comprises an oxidation enhancer and an APOBEC enzyme that is specific for non-oxidized cytosines.

20. The method of any of claims 1-19, wherein the method further comprises denaturing double stranded nucleic acids using heat, a denaturing agent, or a combination thereof.

21. The method of any of claims 1-20, wherein the method further comprises fragmenting nucleic acids and adding an adapter sequence to an end of the nucleic acids prior to contacting nucleic acids with a methylation-specific deamination agent.

22. The method of claim 21, wherein fragmenting nucleic acids and adding an adapter sequence to an end of the nucleic acids comprises tagmenting nucleic acids using a transposome complex.

23. The method of any of claims 1-22, wherein introducing mutations into the deaminated nucleic acids to generate mutated and deaminated nucleic acids comprises amplifying the deaminated nucleic acids with a polymerase, dNTPs, and a nucleotide analog.

24. A kit for determining a sequence of a nucleic acid, the kit comprising: a methylation-specific deamination agent, and one or more mutagenesis agents.

25. A computer-implemented method for determining a sequence of a nucleic acid, the method comprising: receiving first sequence reads from mutated and deaminated nucleic acids; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

26. A system for determining a sequence of a nucleic acid, comprising one or more processors having instructions that when executed perform a method comprising: receiving first sequence reads from mutated and deaminated nucleic acids; and assembling the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

27. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive first sequence reads from mutated and deaminated nucleic acids; and assemble the first sequence reads based on patterns of mutations and deaminations in the mutated and deaminated nucleic acids, thereby determining a nucleic acid sequence.

Citation Information

Patent Citations

  • Method for Introducing Mutations

    US20210010008A1

  • Sequencing Algorithm

    US20210174905A1

  • Spectrum searching method that uses non-chemical qualities of the measurement

    US60635410P0

  • Bisulfite-free, base-resolution identification of cytosine modifications

    WO2019136413A1

  • Sequencing algorithm

    WO2020035669A1